{
  "id": 207741,
  "title": "Welcome from the host",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/207741",
  "author_name": "Ha Q. Nguyen",
  "post_date": "2020-12-31T04:49:19.605000",
  "votes": 63,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>The <a href=\"https://vindr.ai/\" target=\"_blank\">VinDr</a> team at Vingroup Big Data Institute (<a href=\"https://vinbigdata.org/en/\" target=\"_blank\">VinBigdata</a>) is very excited to kick off the challenge on localization and classification of abnormalities from chest X-ray images. Thank you all for participating.</p>\n<p>In this competition, you’re given a set of training X-ray images in DICOM format, each of which was blindly annotated with bounding boxes of 14 classes by 3 radiologists from a pool of 17, encoded with Rad IDs from R1 to R17. The task is to automatically correctly predict boxes around abnormalities and classify them for the test images, whose ground-truth labels are hidden. Unlike the labels of the training set, those of the test set were already a consensus of 5 radiologists per image. </p>\n<p>We have attempted to solve this problem on our own and even deployed the models in our products, but we’re aware that this is still a challenging problem and there is a lot of room for improvement. Through this competition, we would like to see novel ideas in processing DICOM images, harmonizing different opinions of radiologists, training predictive models, and so on.</p>\n<p>The dataset (VinDr-CXR) used in this competition was created through the collaboration between VinBigdata and two hospitals in Vietnam, the Hospital 108 and the Hanoi Medical University Hospital. We would like to make this dataset public since we believe that data sharing is the best way to accelerate the development of machine learning algorithms for medical applications. We encourage you to look at our data descriptor paper for more details. </p>\n<p>We have two Kaggle Masters in the host team, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">Nhan T. Nguyen</a> and <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">Dung B. Nguyen</a>, who will serve as our technical contact points for this competition.      </p>\n<p>Good luck!</p>\n<p>Ha Q. Nguyen<br>\nHead of Medical Imaging Department, VinBigdata</p>",
  "messages": [
    {
      "id": 1133281,
      "postDate": "2020-12-31T04:49:19.607Z",
      "content": "<p>Hi there,</p>\n<p>The <a href=\"https://vindr.ai/\" target=\"_blank\">VinDr</a> team at Vingroup Big Data Institute (<a href=\"https://vinbigdata.org/en/\" target=\"_blank\">VinBigdata</a>) is very excited to kick off the challenge on localization and classification of abnormalities from chest X-ray images. Thank you all for participating.</p>\n<p>In this competition, you’re given a set of training X-ray images in DICOM format, each of which was blindly annotated with bounding boxes of 14 classes by 3 radiologists from a pool of 17, encoded with Rad IDs from R1 to R17. The task is to automatically correctly predict boxes around abnormalities and classify them for the test images, whose ground-truth labels are hidden. Unlike the labels of the training set, those of the test set were already a consensus of 5 radiologists per image. </p>\n<p>We have attempted to solve this problem on our own and even deployed the models in our products, but we’re aware that this is still a challenging problem and there is a lot of room for improvement. Through this competition, we would like to see novel ideas in processing DICOM images, harmonizing different opinions of radiologists, training predictive models, and so on.</p>\n<p>The dataset (VinDr-CXR) used in this competition was created through the collaboration between VinBigdata and two hospitals in Vietnam, the Hospital 108 and the Hanoi Medical University Hospital. We would like to make this dataset public since we believe that data sharing is the best way to accelerate the development of machine learning algorithms for medical applications. We encourage you to look at our data descriptor paper for more details. </p>\n<p>We have two Kaggle Masters in the host team, <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">Nhan T. Nguyen</a> and <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">Dung B. Nguyen</a>, who will serve as our technical contact points for this competition.      </p>\n<p>Good luck!</p>\n<p>Ha Q. Nguyen<br>\nHead of Medical Imaging Department, VinBigdata</p>",
      "rawMarkdown": "Hi there,\n\nThe [VinDr](https://vindr.ai/) team at Vingroup Big Data Institute ([VinBigdata](https://vinbigdata.org/en/)) is very excited to kick off the challenge on localization and classification of abnormalities from chest X-ray images. Thank you all for participating.\n\nIn this competition, you’re given a set of training X-ray images in DICOM format, each of which was blindly annotated with bounding boxes of 14 classes by 3 radiologists from a pool of 17, encoded with Rad IDs from R1 to R17. The task is to automatically correctly predict boxes around abnormalities and classify them for the test images, whose ground-truth labels are hidden. Unlike the labels of the training set, those of the test set were already a consensus of 5 radiologists per image. \n\nWe have attempted to solve this problem on our own and even deployed the models in our products, but we’re aware that this is still a challenging problem and there is a lot of room for improvement. Through this competition, we would like to see novel ideas in processing DICOM images, harmonizing different opinions of radiologists, training predictive models, and so on.\n\nThe dataset (VinDr-CXR) used in this competition was created through the collaboration between VinBigdata and two hospitals in Vietnam, the Hospital 108 and the Hanoi Medical University Hospital. We would like to make this dataset public since we believe that data sharing is the best way to accelerate the development of machine learning algorithms for medical applications. We encourage you to look at our data descriptor paper for more details. \n\nWe have two Kaggle Masters in the host team, [Nhan T. Nguyen](https://www.kaggle.com/andy2709) and [Dung B. Nguyen](https://www.kaggle.com/nguyenbadung), who will serve as our technical contact points for this competition.      \n\nGood luck!\n\nHa Q. Nguyen\nHead of Medical Imaging Department, VinBigdata\n",
      "votes": 62
    },
    {
      "id": 1133286,
      "postDate": "2020-12-31T04:59:11.897Z",
      "content": "<p>I am very proud of VinBigData when organizing this competition. A competition is of great significance not only in terms of medical image processing but also its practical significance of helping doctors a lot not only in Vietnam. Hope to learn more from this competition. Thank you VinBigData for organizing a competition with this large scale and wish the contest will be held successfully and successfully! 💯💕🔥</p>",
      "rawMarkdown": "I am very proud of VinBigData when organizing this competition. A competition is of great significance not only in terms of medical image processing but also its practical significance of helping doctors a lot not only in Vietnam. Hope to learn more from this competition. Thank you VinBigData for organizing a competition with this large scale and wish the contest will be held successfully and successfully! 💯💕🔥",
      "votes": 12
    },
    {
      "id": 1152799,
      "postDate": "2021-01-14T12:58:14.793Z",
      "content": "<p>I would like to get clarity on the Competition metric.</p>\n<p>Since the training set and test set is annotated by Human, I have a confusion</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fe0de1918d70b0981b93d9195e7475a32%2F__results___36_1.png?generation=1610628973015921&amp;alt=media\" alt=\"\"></p>\n<p>You can see here that multiple boxes were linked to same disease and disease space.<br>\nBut my question is will your competition metric calculates True Positive for each of the overlapping boxes</p>\n<p>Thanks in Advance. Just for clarification</p>",
      "rawMarkdown": "I would like to get clarity on the Competition metric.\n\nSince the training set and test set is annotated by Human, I have a confusion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fe0de1918d70b0981b93d9195e7475a32%2F__results___36_1.png?generation=1610628973015921&alt=media)\n\nYou can see here that multiple boxes were linked to same disease and disease space.\nBut my question is will your competition metric calculates True Positive for each of the overlapping boxes\n\nThanks in Advance. Just for clarification",
      "votes": 7,
      "replies": [
        {
          "id": 1157498,
          "postDate": "2021-01-18T00:38:08.883Z",
          "content": "<p>Great question <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> . I'm also interested as well. So for the bottom orange bound box predictions, are there any False Positive penalty for predicting multiple bounding boxes? So is the orange bounding boxes example in the bottom right going to be counted as a True Positive, or 2 False Positives and 1 True positive? <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a></p>\n<p>Thanks for the help,</p>",
          "rawMarkdown": "Great question @morizin . I'm also interested as well. So for the bottom orange bound box predictions, are there any False Positive penalty for predicting multiple bounding boxes? So is the orange bounding boxes example in the bottom right going to be counted as a True Positive, or 2 False Positives and 1 True positive? @nguyenbadung @andy2709\n\nThanks for the help,"
        },
        {
          "id": 1158212,
          "postDate": "2021-01-18T13:02:54.697Z",
          "content": "<p>We are expecting the host responsive please <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> </p>",
          "rawMarkdown": "We are expecting the host responsive please @nguyenbadung @andy2709 @nguyenquyha "
        },
        {
          "id": 1159126,
          "postDate": "2021-01-19T04:14:23.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> The reason why multiple boxes were linked to same disease was that each image in the train data set was independently annotated by 3 different rads. Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes. In my opinion, you need to think of creative ways to combine our provided bounding boxes in the train data set before conducting any experiment.</p>\n<p>The competition metric is a variant of the PascalVOC, you can refer to this document <a href=\"http://host.robots.ox.ac.uk/pascal/VOC/voc2010/devkit_doc_08-May-2010.pdf\" target=\"_blank\">http://host.robots.ox.ac.uk/pascal/VOC/voc2010/devkit_doc_08-May-2010.pdf</a> for further details. In section 4.4 on page 11, it says:</p>\n<pre><code>Example code for computing this overlap measure is provided in the development kit. Multiple detections of the same object in an image are considered\nfalse detections e.g. 5 detections of a single object is counted as 1 correct detection and 4 false detections – it is the responsibility of the participant’s system\nto filter multiple detections from its output.\n</code></pre>\n<p>For predicting the test set, you <strong>have to apply</strong> NMS to filter multiple detections of the same lesion region.</p>\n<p>Also, these two discussion threads by Peter and Phalanx are very helpful<br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212287\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212287</a><br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/208837\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/208837</a></p>",
          "rawMarkdown": "@morizin The reason why multiple boxes were linked to same disease was that each image in the train data set was independently annotated by 3 different rads. Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes. In my opinion, you need to think of creative ways to combine our provided bounding boxes in the train data set before conducting any experiment.\n\nThe competition metric is a variant of the PascalVOC, you can refer to this document http://host.robots.ox.ac.uk/pascal/VOC/voc2010/devkit_doc_08-May-2010.pdf for further details. In section 4.4 on page 11, it says:\n```\nExample code for computing this overlap measure is provided in the development kit. Multiple detections of the same object in an image are considered\nfalse detections e.g. 5 detections of a single object is counted as 1 correct detection and 4 false detections – it is the responsibility of the participant’s system\nto filter multiple detections from its output.\n```\nFor predicting the test set, you **have to apply** NMS to filter multiple detections of the same lesion region.\n\nAlso, these two discussion threads by Peter and Phalanx are very helpful\nhttps://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212287\nhttps://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/208837\n",
          "votes": 14
        },
        {
          "id": 1159998,
          "postDate": "2021-01-19T15:48:33.417Z",
          "content": "<p>Thank you for great clarification <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
          "rawMarkdown": "Thank you for great clarification @andy2709 "
        },
        {
          "id": 1163896,
          "postDate": "2021-01-22T03:11:57.323Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1170601,
          "postDate": "2021-01-26T10:08:54.370Z",
          "content": "<p>If there are multiple detections of the same object, which one your implementation selects as correct? The one with largest IoU or confidence? <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
          "rawMarkdown": "If there are multiple detections of the same object, which one your implementation selects as correct? The one with largest IoU or confidence? @andy2709 "
        },
        {
          "id": 1175808,
          "postDate": "2021-01-29T11:05:53.930Z",
          "content": "<p>In original Pascal VOC implementation predictions are sorted by confidence before computing the average precision, so in case of multiple detections only the one with largest confidence is considered as TP. I think it's the same here, but a clarification from the host would be appreciated.</p>",
          "rawMarkdown": "In original Pascal VOC implementation predictions are sorted by confidence before computing the average precision, so in case of multiple detections only the one with largest confidence is considered as TP. I think it's the same here, but a clarification from the host would be appreciated.",
          "votes": 2
        },
        {
          "id": 1216738,
          "postDate": "2021-02-24T13:02:52.277Z",
          "content": "<blockquote>\n  <p>Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> did they add or subtract any bounding boxes while 'unify' the annotations<br>\nor they will just merge the annotations identified by the previous 3 rads</p>",
          "rawMarkdown": "> Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes.\n\n@andy2709 did they add or subtract any bounding boxes while 'unify' the annotations\nor they will just merge the annotations identified by the previous 3 rads",
          "votes": 1
        }
      ]
    },
    {
      "id": 1133975,
      "postDate": "2020-12-31T17:46:20.820Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> thank you for organizing this competition!</p>\n<p>The general rules indicate:</p>\n<blockquote>\n  <p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n</blockquote>\n<p>I'd like to create a preprocessing <strong>dataset</strong> to help participants more easily load the data, and I will of course link to this competition and its rules. <strong>Is that be allowed?</strong></p>",
      "rawMarkdown": "@nguyenquyha thank you for organizing this competition!\n\nThe general rules indicate:\n> B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nI'd like to create a preprocessing **dataset** to help participants more easily load the data, and I will of course link to this competition and its rules. **Is that be allowed?**",
      "votes": 5,
      "replies": [
        {
          "id": 1135238,
          "postDate": "2021-01-02T03:56:38.230Z",
          "content": "<p>We're totally fine with that, as long as you point the derived dataset to this competition.</p>",
          "rawMarkdown": "We're totally fine with that, as long as you point the derived dataset to this competition.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1153188,
      "postDate": "2021-01-14T17:01:14.693Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> Thank you for hosting this competition.<br>\nI'm interested in this task.</p>\n<p>I have a question about labels.</p>\n<p>radiologist cannot see another rad_id labeling?<br>\nif the radiologist check rad0, the annotated image in rad1,2 can not be seen by rad0 radiologist</p>",
      "rawMarkdown": "@nguyenquyha Thank you for hosting this competition.\nI'm interested in this task.\n\nI have a question about labels.\n\nradiologist cannot see another rad_id labeling?\nif the radiologist check rad0, the annotated image in rad1,2 can not be seen by rad0 radiologist",
      "votes": 3,
      "replies": [
        {
          "id": 1153195,
          "postDate": "2021-01-14T17:07:47.583Z",
          "content": "<p>Maybe sometimes the rad0 maybe could miss it. they are humans</p>",
          "rawMarkdown": "Maybe sometimes the rad0 maybe could miss it. they are humans"
        },
        {
          "id": 1153490,
          "postDate": "2021-01-14T23:50:52.043Z",
          "content": "<p>I agree. the labeling process is very important for us.<br>\nI want to clearly </p>",
          "rawMarkdown": "I agree. the labeling process is very important for us.\nI want to clearly "
        },
        {
          "id": 1160693,
          "postDate": "2021-01-20T04:33:35.087Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nIn trainset, the radiologists didn't see each other's labels</p>",
          "rawMarkdown": "Hi @tereka \nIn trainset, the radiologists didn't see each other's labels",
          "votes": 3
        },
        {
          "id": 1160703,
          "postDate": "2021-01-20T04:40:47.880Z",
          "content": "<p>Thank you for the clarification</p>",
          "rawMarkdown": "Thank you for the clarification"
        },
        {
          "id": 1167100,
          "postDate": "2021-01-24T03:47:48.533Z",
          "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <br>\nthank you for the clarification!</p>",
          "rawMarkdown": "@nguyenbadung \nthank you for the clarification!"
        }
      ]
    },
    {
      "id": 1232780,
      "postDate": "2021-03-10T01:50:46.460Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> </p>\n<p>Quick question regarding the public test set distribution.</p>\n<p>Should we expect the label distribution ( various abnormalities &amp; normal ) to be similar to the private set ?   Or it could be quite different ? </p>",
      "rawMarkdown": "@nguyenquyha \n\nQuick question regarding the public test set distribution.\n\nShould we expect the label distribution ( various abnormalities & normal ) to be similar to the private set ?   Or it could be quite different ? ",
      "votes": 1
    },
    {
      "id": 1209998,
      "postDate": "2021-02-19T06:37:28.517Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a><br>\nCan we use external dataset<br>\nI mean of Chexpert , MIMIC ,NIH chest14.<br>\nCan we use the data of ongoing competition Ranzcr.</p>",
      "rawMarkdown": "Hi @nguyenbadung @andy2709 @nguyenquyha\nCan we use external dataset\nI mean of Chexpert , MIMIC ,NIH chest14.\nCan we use the data of ongoing competition Ranzcr.",
      "votes": 1,
      "replies": [
        {
          "id": 1230379,
          "postDate": "2021-03-08T04:53:39.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a><br>\nto get access to Chexpert and MIMIC data, we need to be a PhysioNet credentialing user.<br>\nseems to require some training to become one.<br>\nIs such data still considered as public data?</p>",
          "rawMarkdown": "@nguyenbadung @andy2709 @nguyenquyha\nto get access to Chexpert and MIMIC data, we need to be a PhysioNet credentialing user.\nseems to require some training to become one.\nIs such data still considered as public data?",
          "votes": 1
        },
        {
          "id": 1230399,
          "postDate": "2021-03-08T05:09:14.567Z",
          "content": "<p>Yes, we consider CheXpert, MIMIC as public datasets.</p>",
          "rawMarkdown": "Yes, we consider CheXpert, MIMIC as public datasets.",
          "votes": 1
        },
        {
          "id": 1230765,
          "postDate": "2021-03-08T12:50:01.847Z",
          "content": "<p>Thanks for clarification</p>",
          "rawMarkdown": "Thanks for clarification"
        }
      ]
    },
    {
      "id": 1157848,
      "postDate": "2021-01-18T07:52:34.077Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2F8bafe485431584320ede6c5df26062c9%2F__results___61_0.png?generation=1610956184045040&amp;alt=media\" alt=\"\"></p>\n<p>Hi <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>,Nhan and Dung, can you please help me with the bottom left image. There are big and small overlapping bounding boxes for the class <code>Nodule/Mass</code> for the image in the lower left. How do we make sense of this? Can you give some insight, because it seems like maybe one annotator maybe just annotated two big boxes on the left and right lung for 'Nodule/Mass' while another annotated many small boxes.</p>\n<p>Thanks for the help, </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2F8bafe485431584320ede6c5df26062c9%2F__results___61_0.png?generation=1610956184045040&alt=media)\n\nHi @nguyenbadung @andy2709,Nhan and Dung, can you please help me with the bottom left image. There are big and small overlapping bounding boxes for the class `Nodule/Mass` for the image in the lower left. How do we make sense of this? Can you give some insight, because it seems like maybe one annotator maybe just annotated two big boxes on the left and right lung for 'Nodule/Mass' while another annotated many small boxes.\n\nThanks for the help, ",
      "votes": 1,
      "replies": [
        {
          "id": 1157987,
          "postDate": "2021-01-18T10:13:50.770Z",
          "content": "<p>Yes, you're right, It seems like the annotator just draws one big box for a group of abnormalities.<br>\nIt's human problem and we can't control it.</p>",
          "rawMarkdown": "Yes, you're right, It seems like the annotator just draws one big box for a group of abnormalities.\nIt's human problem and we can't control it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1147819,
      "postDate": "2021-01-10T18:07:39.030Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> <br>\nThanks for your work!<br>\nMay I ask you about the algorithm of consensus for the test data? Was it just a result of the discussion of 5 radiologists in one room or you received predictions from each radiologist and went this data through some consensus algorithm?</p>",
      "rawMarkdown": "@nguyenquyha \nThanks for your work!\nMay I ask you about the algorithm of consensus for the test data? Was it just a result of the discussion of 5 radiologists in one room or you received predictions from each radiologist and went this data through some consensus algorithm?",
      "votes": 1,
      "replies": [
        {
          "id": 1148226,
          "postDate": "2021-01-11T02:44:04.680Z",
          "content": "<p>It was mentioned in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">our paper</a>:</p>\n<p>\"For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\"</p>",
          "rawMarkdown": "It was mentioned in [our paper](https://arxiv.org/pdf/2012.15029.pdf):\n\n\"For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\"",
          "votes": 3
        },
        {
          "id": 1191687,
          "postDate": "2021-02-08T15:55:05.900Z",
          "content": "<p>Thanks for sharing this amazing dataset. Can you tell us a little bit how multiple correctly labeled bounding boxes (size, location) were average by the 2 experienced annotators ? Let say all 3 initial annotators correctly labeled a nodule with small difference in size and location. Did the 2 experienced radiologists just selected 1 of the 3 bounding boxes as a final label or there was some kind of averaging ? Thx a lot ! </p>",
          "rawMarkdown": "Thanks for sharing this amazing dataset. Can you tell us a little bit how multiple correctly labeled bounding boxes (size, location) were average by the 2 experienced annotators ? Let say all 3 initial annotators correctly labeled a nodule with small difference in size and location. Did the 2 experienced radiologists just selected 1 of the 3 bounding boxes as a final label or there was some kind of averaging ? Thx a lot ! "
        }
      ]
    },
    {
      "id": 1135443,
      "postDate": "2021-01-02T08:56:54.577Z",
      "content": "<p>Thanks for organizing this competition! <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> . I have few questions below:</p>\n<ol>\n<li>Is the full (public + private) test just containing 3000 images? Or 3000 images are just used for public LB? (IMO, to prevent the issue of just forking and blending, the private data should be hidden).</li>\n<li>I haven't seen the rule of this competition, e.g. Are we allowed to use pre-trained models or external data?</li>\n</ol>",
      "rawMarkdown": "Thanks for organizing this competition! @nguyenquyha . I have few questions below:\n1. Is the full (public + private) test just containing 3000 images? Or 3000 images are just used for public LB? (IMO, to prevent the issue of just forking and blending, the private data should be hidden).\n2. I haven't seen the rule of this competition, e.g. Are we allowed to use pre-trained models or external data?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1135970,
          "postDate": "2021-01-02T16:28:40.337Z",
          "content": "<ol>\n<li><p>The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB</p></li>\n<li><p>Yes, pre-trained models and external data are allowed provided that they meet the rules stated in Section 7.C. In general, repositories like detectron2, mmdet, torchvision, timm, efficientdet with pre-trained ImageNet/COCO/OpenImages models and datasets like Chexpert, MIMIC, PadChest <strong>are allowed</strong></p></li>\n</ol>",
          "rawMarkdown": "1. The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB\n\n2. Yes, pre-trained models and external data are allowed provided that they meet the rules stated in Section 7.C. In general, repositories like detectron2, mmdet, torchvision, timm, efficientdet with pre-trained ImageNet/COCO/OpenImages models and datasets like Chexpert, MIMIC, PadChest **are allowed**",
          "votes": 3
        },
        {
          "id": 1142015,
          "postDate": "2021-01-07T04:36:39.413Z",
          "content": "<p>YOLO v5 which is  GPL-3.0 Licensed; can we use that or only MIT licensed ones? I am getting confused</p>",
          "rawMarkdown": "YOLO v5 which is  GPL-3.0 Licensed; can we use that or only MIT licensed ones? I am getting confused"
        },
        {
          "id": 1142016,
          "postDate": "2021-01-07T04:40:29.463Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207679\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207679</a></p>",
          "rawMarkdown": "https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207679",
          "votes": 1
        }
      ]
    },
    {
      "id": 1133631,
      "postDate": "2020-12-31T11:46:12.930Z",
      "content": "<p>Thanks for the host and for the amazing dataset. </p>\n<p>Knowing some great kagglers are among organizers will surely help for quick and accurate feedback :)</p>",
      "rawMarkdown": "Thanks for the host and for the amazing dataset. \n\nKnowing some great kagglers are among organizers will surely help for quick and accurate feedback :)",
      "votes": 1
    },
    {
      "id": 1233551,
      "postDate": "2021-03-10T13:57:08.557Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> <br>\nThanks for hosting this interesting competition.<br>\nI am wondering if semi-supervised learning methods such as pseudo-labeling is allowed in this competition? <br>\nI believe it should be allowed, but I still want a confirmation from host. </p>",
      "rawMarkdown": "@nguyenbadung @andy2709 @nguyenquyha \nThanks for hosting this interesting competition.\nI am wondering if semi-supervised learning methods such as pseudo-labeling is allowed in this competition? \nI believe it should be allowed, but I still want a confirmation from host. ",
      "votes": 2
    },
    {
      "id": 1134118,
      "postDate": "2020-12-31T21:11:39.667Z",
      "content": "<p>This is really an interesting problem and thanks for hosting this competition. </p>",
      "rawMarkdown": "This is really an interesting problem and thanks for hosting this competition. ",
      "votes": 2
    },
    {
      "id": 1563614,
      "postDate": "2021-10-28T13:06:17.040Z",
      "content": "<p>Hello dear<br>\nWhy MRI images are not used instead of X-rays to diagnose lung disease? <br>\nAren't MRI images more accurate for this?</p>",
      "rawMarkdown": "Hello dear\nWhy MRI images are not used instead of X-rays to diagnose lung disease? \nAren't MRI images more accurate for this?"
    },
    {
      "id": 1462434,
      "postDate": "2021-08-09T20:42:39.013Z",
      "content": "<p>Hi ! I am new to Kaggle and trying to explore this competition and i am little confused about the values mentioned under Score within Leaderboard.<br>\nSo, are these MAP (Mean average precision) values of each model? or is this something else?<br>\nthank you.</p>",
      "rawMarkdown": "Hi ! I am new to Kaggle and trying to explore this competition and i am little confused about the values mentioned under Score within Leaderboard.\nSo, are these MAP (Mean average precision) values of each model? or is this something else?\nthank you."
    },
    {
      "id": 1457267,
      "postDate": "2021-08-07T10:27:04.360Z",
      "content": "<p>i agree with you, kind sir</p>",
      "rawMarkdown": "i agree with you, kind sir\n"
    },
    {
      "id": 1256347,
      "postDate": "2021-03-29T19:04:38.817Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> , <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> , <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> ,<br>\nnow, that the end of the competition is near, i would like to ask you something technical.<br>\nI found some attenuation patterns on many of the images: namely hyperlucent hemithoraces, mainly on the left side like here :<br>\n<img src=\"https://i.postimg.cc/RhDdzZJ8/attenuation.jpg\" alt=\"att\"><br>\nThese are not oblique chest radiographs - that could explain this phenomenon -, we can furthermore exclude the other obvious causes (absence of a breast after mastectomy, absence of a pectoralis muscle, unilateral or asymmetric bullous disease/emphysema, or air trapping, or an endobronchial foreign body). It made me really hard to evaluate the lung parenchyma!</p>\n<p>Is that a general technical issue? I must admit i do not encounter such images in our daily routine.</p>\n<p>Thank you in advance.</p>",
      "rawMarkdown": "@nguyenquyha , @andy2709 , @nguyenbadung ,\nnow, that the end of the competition is near, i would like to ask you something technical.\nI found some attenuation patterns on many of the images: namely hyperlucent hemithoraces, mainly on the left side like here :\n![att](https://i.postimg.cc/RhDdzZJ8/attenuation.jpg)\nThese are not oblique chest radiographs - that could explain this phenomenon -, we can furthermore exclude the other obvious causes (absence of a breast after mastectomy, absence of a pectoralis muscle, unilateral or asymmetric bullous disease/emphysema, or air trapping, or an endobronchial foreign body). It made me really hard to evaluate the lung parenchyma!\n\nIs that a general technical issue? I must admit i do not encounter such images in our daily routine.\n\nThank you in advance.\n\n\n"
    },
    {
      "id": 1192268,
      "postDate": "2021-02-09T04:14:00.177Z",
      "content": "<p>Hi Ha,<br>\nI'm wondering how can I access the age of the cases?<br>\nThank you.</p>",
      "rawMarkdown": "Hi Ha,\nI'm wondering how can I access the age of the cases?\nThank you."
    },
    {
      "id": 1145500,
      "postDate": "2021-01-09T07:32:35.217Z",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> Thanks for organizing interesting challenge!<br>\nI may miss to read some description, but may I ask how train/test set were split? </p>",
      "rawMarkdown": "@nguyenquyha Thanks for organizing interesting challenge!\nI may miss to read some description, but may I ask how train/test set were split? ",
      "replies": [
        {
          "id": 1148230,
          "postDate": "2021-01-11T02:46:25.217Z",
          "content": "<p>Please refer to Figure 1 in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">our paper</a>.</p>",
          "rawMarkdown": "Please refer to Figure 1 in [our paper](https://arxiv.org/pdf/2012.15029.pdf)."
        }
      ]
    },
    {
      "id": 1732595,
      "postDate": "2022-03-23T14:43:25.777Z",
      "content": "<p>Thanks for clarification.</p>",
      "rawMarkdown": "Thanks for clarification."
    },
    {
      "id": 1143144,
      "postDate": "2021-01-07T19:15:14.033Z",
      "content": "<p>Thanks for organizing this competition! </p>",
      "rawMarkdown": "Thanks for organizing this competition! "
    }
  ],
  "comments": [
    {
      "id": 1133286,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-12-31T04:59:11.897000",
      "content": "<p>I am very proud of VinBigData when organizing this competition. A competition is of great significance not only in terms of medical image processing but also its practical significance of helping doctors a lot not only in Vietnam. Hope to learn more from this competition. Thank you VinBigData for organizing a competition with this large scale and wish the contest will be held successfully and successfully! 💯💕🔥</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 1152799,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-01-14T12:58:14.793000",
      "content": "<p>I would like to get clarity on the Competition metric.</p>\n<p>Since the training set and test set is annotated by Human, I have a confusion</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fe0de1918d70b0981b93d9195e7475a32%2F__results___36_1.png?generation=1610628973015921&amp;alt=media\" alt=\"\"></p>\n<p>You can see here that multiple boxes were linked to same disease and disease space.<br>\nBut my question is will your competition metric calculates True Positive for each of the overlapping boxes</p>\n<p>Thanks in Advance. Just for clarification</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1157498,
          "author_name": "Bo Peng",
          "author_url": "",
          "post_date": "2021-01-18T00:38:08.883000",
          "content": "<p>Great question <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> . I'm also interested as well. So for the bottom orange bound box predictions, are there any False Positive penalty for predicting multiple bounding boxes? So is the orange bounding boxes example in the bottom right going to be counted as a True Positive, or 2 False Positives and 1 True positive? <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a></p>\n<p>Thanks for the help,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1158212,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-18T13:02:54.697000",
          "content": "<p>We are expecting the host responsive please <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1159126,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2021-01-19T04:14:23.063000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> The reason why multiple boxes were linked to same disease was that each image in the train data set was independently annotated by 3 different rads. Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes. In my opinion, you need to think of creative ways to combine our provided bounding boxes in the train data set before conducting any experiment.</p>\n<p>The competition metric is a variant of the PascalVOC, you can refer to this document <a href=\"http://host.robots.ox.ac.uk/pascal/VOC/voc2010/devkit_doc_08-May-2010.pdf\" target=\"_blank\">http://host.robots.ox.ac.uk/pascal/VOC/voc2010/devkit_doc_08-May-2010.pdf</a> for further details. In section 4.4 on page 11, it says:</p>\n<pre><code>Example code for computing this overlap measure is provided in the development kit. Multiple detections of the same object in an image are considered\nfalse detections e.g. 5 detections of a single object is counted as 1 correct detection and 4 false detections – it is the responsibility of the participant’s system\nto filter multiple detections from its output.\n</code></pre>\n<p>For predicting the test set, you <strong>have to apply</strong> NMS to filter multiple detections of the same lesion region.</p>\n<p>Also, these two discussion threads by Peter and Phalanx are very helpful<br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212287\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212287</a><br>\n<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/208837\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/208837</a></p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1159998,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-19T15:48:33.417000",
          "content": "<p>Thank you for great clarification <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1163896,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-22T03:11:57.323000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1170601,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-01-26T10:08:54.370000",
          "content": "<p>If there are multiple detections of the same object, which one your implementation selects as correct? The one with largest IoU or confidence? <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1175808,
          "author_name": "LAZCoder",
          "author_url": "",
          "post_date": "2021-01-29T11:05:53.930000",
          "content": "<p>In original Pascal VOC implementation predictions are sorted by confidence before computing the average precision, so in case of multiple detections only the one with largest confidence is considered as TP. I think it's the same here, but a clarification from the host would be appreciated.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1216738,
          "author_name": "Shrijeet16",
          "author_url": "",
          "post_date": "2021-02-24T13:02:52.277000",
          "content": "<blockquote>\n  <p>Such cases won't exist in the test data set since we asked 2 more rads to review and 'unify' all the bounding boxes.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> did they add or subtract any bounding boxes while 'unify' the annotations<br>\nor they will just merge the annotations identified by the previous 3 rads</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1133975,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2020-12-31T17:46:20.820000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> thank you for organizing this competition!</p>\n<p>The general rules indicate:</p>\n<blockquote>\n  <p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n</blockquote>\n<p>I'd like to create a preprocessing <strong>dataset</strong> to help participants more easily load the data, and I will of course link to this competition and its rules. <strong>Is that be allowed?</strong></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1135238,
          "author_name": "Ha Q. Nguyen",
          "author_url": "",
          "post_date": "2021-01-02T03:56:38.230000",
          "content": "<p>We're totally fine with that, as long as you point the derived dataset to this competition.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1153188,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2021-01-14T17:01:14.693000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> Thank you for hosting this competition.<br>\nI'm interested in this task.</p>\n<p>I have a question about labels.</p>\n<p>radiologist cannot see another rad_id labeling?<br>\nif the radiologist check rad0, the annotated image in rad1,2 can not be seen by rad0 radiologist</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1153195,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-14T17:07:47.583000",
          "content": "<p>Maybe sometimes the rad0 maybe could miss it. they are humans</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1153490,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2021-01-14T23:50:52.043000",
          "content": "<p>I agree. the labeling process is very important for us.<br>\nI want to clearly </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1160693,
          "author_name": "DungNB",
          "author_url": "",
          "post_date": "2021-01-20T04:33:35.087000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> <br>\nIn trainset, the radiologists didn't see each other's labels</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1160703,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-20T04:40:47.880000",
          "content": "<p>Thank you for the clarification</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1167100,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2021-01-24T03:47:48.533000",
          "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <br>\nthank you for the clarification!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1232780,
      "author_name": "NakedKoala",
      "author_url": "",
      "post_date": "2021-03-10T01:50:46.460000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> </p>\n<p>Quick question regarding the public test set distribution.</p>\n<p>Should we expect the label distribution ( various abnormalities &amp; normal ) to be similar to the private set ?   Or it could be quite different ? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1209998,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-02-19T06:37:28.517000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a><br>\nCan we use external dataset<br>\nI mean of Chexpert , MIMIC ,NIH chest14.<br>\nCan we use the data of ongoing competition Ranzcr.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1230379,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2021-03-08T04:53:39.127000",
          "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a><br>\nto get access to Chexpert and MIMIC data, we need to be a PhysioNet credentialing user.<br>\nseems to require some training to become one.<br>\nIs such data still considered as public data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1230399,
          "author_name": "Ha Q. Nguyen",
          "author_url": "",
          "post_date": "2021-03-08T05:09:14.567000",
          "content": "<p>Yes, we consider CheXpert, MIMIC as public datasets.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1230765,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-08T12:50:01.847000",
          "content": "<p>Thanks for clarification</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1157848,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2021-01-18T07:52:34.077000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2F8bafe485431584320ede6c5df26062c9%2F__results___61_0.png?generation=1610956184045040&amp;alt=media\" alt=\"\"></p>\n<p>Hi <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a>,Nhan and Dung, can you please help me with the bottom left image. There are big and small overlapping bounding boxes for the class <code>Nodule/Mass</code> for the image in the lower left. How do we make sense of this? Can you give some insight, because it seems like maybe one annotator maybe just annotated two big boxes on the left and right lung for 'Nodule/Mass' while another annotated many small boxes.</p>\n<p>Thanks for the help, </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1157987,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2021-01-18T10:13:50.770000",
          "content": "<p>Yes, you're right, It seems like the annotator just draws one big box for a group of abnormalities.<br>\nIt's human problem and we can't control it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1147819,
      "author_name": "Leonid Verkhovtsev",
      "author_url": "",
      "post_date": "2021-01-10T18:07:39.030000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> <br>\nThanks for your work!<br>\nMay I ask you about the algorithm of consensus for the test data? Was it just a result of the discussion of 5 radiologists in one room or you received predictions from each radiologist and went this data through some consensus algorithm?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1148226,
          "author_name": "Ha Q. Nguyen",
          "author_url": "",
          "post_date": "2021-01-11T02:44:04.680000",
          "content": "<p>It was mentioned in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">our paper</a>:</p>\n<p>\"For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\"</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1191687,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2021-02-08T15:55:05.900000",
          "content": "<p>Thanks for sharing this amazing dataset. Can you tell us a little bit how multiple correctly labeled bounding boxes (size, location) were average by the 2 experienced annotators ? Let say all 3 initial annotators correctly labeled a nodule with small difference in size and location. Did the 2 experienced radiologists just selected 1 of the 3 bounding boxes as a final label or there was some kind of averaging ? Thx a lot ! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1135443,
      "author_name": "The fearless",
      "author_url": "",
      "post_date": "2021-01-02T08:56:54.577000",
      "content": "<p>Thanks for organizing this competition! <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> . I have few questions below:</p>\n<ol>\n<li>Is the full (public + private) test just containing 3000 images? Or 3000 images are just used for public LB? (IMO, to prevent the issue of just forking and blending, the private data should be hidden).</li>\n<li>I haven't seen the rule of this competition, e.g. Are we allowed to use pre-trained models or external data?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1135970,
          "author_name": "NguyenThanhNhan",
          "author_url": "",
          "post_date": "2021-01-02T16:28:40.337000",
          "content": "<ol>\n<li><p>The test set contains 3000 images. The public score will be calculated on approx. 10% of the test set and will be visible to you during the competition. Final rankings will be determined on the remaining 90% i.e. private LB</p></li>\n<li><p>Yes, pre-trained models and external data are allowed provided that they meet the rules stated in Section 7.C. In general, repositories like detectron2, mmdet, torchvision, timm, efficientdet with pre-trained ImageNet/COCO/OpenImages models and datasets like Chexpert, MIMIC, PadChest <strong>are allowed</strong></p></li>\n</ol>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1142015,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2021-01-07T04:36:39.413000",
          "content": "<p>YOLO v5 which is  GPL-3.0 Licensed; can we use that or only MIT licensed ones? I am getting confused</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1142016,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-01-07T04:40:29.463000",
          "content": "<p><a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207679\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207679</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1133631,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-12-31T11:46:12.930000",
      "content": "<p>Thanks for the host and for the amazing dataset. </p>\n<p>Knowing some great kagglers are among organizers will surely help for quick and accurate feedback :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1233551,
      "author_name": "nvnn",
      "author_url": "",
      "post_date": "2021-03-10T13:57:08.557000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> <a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> <br>\nThanks for hosting this interesting competition.<br>\nI am wondering if semi-supervised learning methods such as pseudo-labeling is allowed in this competition? <br>\nI believe it should be allowed, but I still want a confirmation from host. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1134118,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-12-31T21:11:39.667000",
      "content": "<p>This is really an interesting problem and thanks for hosting this competition. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1563614,
      "author_name": "HamidReza69GH",
      "author_url": "",
      "post_date": "2021-10-28T13:06:17.040000",
      "content": "<p>Hello dear<br>\nWhy MRI images are not used instead of X-rays to diagnose lung disease? <br>\nAren't MRI images more accurate for this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1462434,
      "author_name": "Yasir Irfan",
      "author_url": "",
      "post_date": "2021-08-09T20:42:39.013000",
      "content": "<p>Hi ! I am new to Kaggle and trying to explore this competition and i am little confused about the values mentioned under Score within Leaderboard.<br>\nSo, are these MAP (Mean average precision) values of each model? or is this something else?<br>\nthank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1457267,
      "author_name": "Aditya Singh",
      "author_url": "",
      "post_date": "2021-08-07T10:27:04.360000",
      "content": "<p>i agree with you, kind sir</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1256347,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2021-03-29T19:04:38.817000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> , <a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> , <a href=\"https://www.kaggle.com/nguyenbadung\" target=\"_blank\">@nguyenbadung</a> ,<br>\nnow, that the end of the competition is near, i would like to ask you something technical.<br>\nI found some attenuation patterns on many of the images: namely hyperlucent hemithoraces, mainly on the left side like here :<br>\n<img src=\"https://i.postimg.cc/RhDdzZJ8/attenuation.jpg\" alt=\"att\"><br>\nThese are not oblique chest radiographs - that could explain this phenomenon -, we can furthermore exclude the other obvious causes (absence of a breast after mastectomy, absence of a pectoralis muscle, unilateral or asymmetric bullous disease/emphysema, or air trapping, or an endobronchial foreign body). It made me really hard to evaluate the lung parenchyma!</p>\n<p>Is that a general technical issue? I must admit i do not encounter such images in our daily routine.</p>\n<p>Thank you in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1192268,
      "author_name": "ServehPadash",
      "author_url": "",
      "post_date": "2021-02-09T04:14:00.177000",
      "content": "<p>Hi Ha,<br>\nI'm wondering how can I access the age of the cases?<br>\nThank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1145500,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2021-01-09T07:32:35.217000",
      "content": "<p><a href=\"https://www.kaggle.com/nguyenquyha\" target=\"_blank\">@nguyenquyha</a> Thanks for organizing interesting challenge!<br>\nI may miss to read some description, but may I ask how train/test set were split? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148230,
          "author_name": "Ha Q. Nguyen",
          "author_url": "",
          "post_date": "2021-01-11T02:46:25.217000",
          "content": "<p>Please refer to Figure 1 in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">our paper</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1732595,
      "author_name": "Ruixuan ALEX",
      "author_url": "",
      "post_date": "2022-03-23T14:43:25.777000",
      "content": "<p>Thanks for clarification.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143144,
      "author_name": "Tilek",
      "author_url": "",
      "post_date": "2021-01-07T19:15:14.033000",
      "content": "<p>Thanks for organizing this competition! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1133281": "Hi there,\n\nThe [VinDr](https://vindr.ai/) team at Vingroup Big Data Institute ([VinBigdata](https://vinbigdata.org/en/)) is very excited to kick off the challenge on localization and classification of abnormalities from chest X-ray images. Thank you all for participating.\n\nIn this competition, you’re given a set of training X-ray images in DICOM format, each of which was blindly annotated with bounding boxes of 14 classes by 3 radiologists from a pool of 17, encoded with Rad IDs from R1 to R17. The task is to automatically correctly predict boxes around abnormalities and classify them for the test images, whose ground-truth labels are hidden. Unlike the labels of the training set, those of the test set were already a consensus of 5 radiologists per image. \n\nWe have attempted to solve this problem on our own and even deployed the models in our products, but we’re aware that this is still a challenging problem and there is a lot of room for improvement. Through this competition, we would like to see novel ideas in processing DICOM images, harmonizing different opinions of radiologists, training predictive models, and so on.\n\nThe dataset (VinDr-CXR) used in this competition was created through the collaboration between VinBigdata and two hospitals in Vietnam, the Hospital 108 and the Hanoi Medical University Hospital. We would like to make this dataset public since we believe that data sharing is the best way to accelerate the development of machine learning algorithms for medical applications. We encourage you to look at our data descriptor paper for more details. \n\nWe have two Kaggle Masters in the host team, [Nhan T. Nguyen](https://www.kaggle.com/andy2709) and [Dung B. Nguyen](https://www.kaggle.com/nguyenbadung), who will serve as our technical contact points for this competition.      \n\nGood luck!\n\nHa Q. Nguyen\nHead of Medical Imaging Department, VinBigdata\n",
    "1133286": "I am very proud of VinBigData when organizing this competition. A competition is of great significance not only in terms of medical image processing but also its practical significance of helping doctors a lot not only in Vietnam. Hope to learn more from this competition. Thank you VinBigData for organizing a competition with this large scale and wish the contest will be held successfully and successfully! 💯💕🔥",
    "1152799": "I would like to get clarity on the Competition metric.\n\nSince the training set and test set is annotated by Human, I have a confusion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3514793%2Fe0de1918d70b0981b93d9195e7475a32%2F__results___36_1.png?generation=1610628973015921&alt=media)\n\nYou can see here that multiple boxes were linked to same disease and disease space.\nBut my question is will your competition metric calculates True Positive for each of the overlapping boxes\n\nThanks in Advance. Just for clarification",
    "1133975": "@nguyenquyha thank you for organizing this competition!\n\nThe general rules indicate:\n> B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nI'd like to create a preprocessing **dataset** to help participants more easily load the data, and I will of course link to this competition and its rules. **Is that be allowed?**",
    "1153188": "@nguyenquyha Thank you for hosting this competition.\nI'm interested in this task.\n\nI have a question about labels.\n\nradiologist cannot see another rad_id labeling?\nif the radiologist check rad0, the annotated image in rad1,2 can not be seen by rad0 radiologist",
    "1232780": "@nguyenquyha \n\nQuick question regarding the public test set distribution.\n\nShould we expect the label distribution ( various abnormalities & normal ) to be similar to the private set ?   Or it could be quite different ? ",
    "1209998": "Hi @nguyenbadung @andy2709 @nguyenquyha\nCan we use external dataset\nI mean of Chexpert , MIMIC ,NIH chest14.\nCan we use the data of ongoing competition Ranzcr.",
    "1157848": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F385905%2F8bafe485431584320ede6c5df26062c9%2F__results___61_0.png?generation=1610956184045040&alt=media)\n\nHi @nguyenbadung @andy2709,Nhan and Dung, can you please help me with the bottom left image. There are big and small overlapping bounding boxes for the class `Nodule/Mass` for the image in the lower left. How do we make sense of this? Can you give some insight, because it seems like maybe one annotator maybe just annotated two big boxes on the left and right lung for 'Nodule/Mass' while another annotated many small boxes.\n\nThanks for the help, ",
    "1147819": "@nguyenquyha \nThanks for your work!\nMay I ask you about the algorithm of consensus for the test data? Was it just a result of the discussion of 5 radiologists in one room or you received predictions from each radiologist and went this data through some consensus algorithm?",
    "1135443": "Thanks for organizing this competition! @nguyenquyha . I have few questions below:\n1. Is the full (public + private) test just containing 3000 images? Or 3000 images are just used for public LB? (IMO, to prevent the issue of just forking and blending, the private data should be hidden).\n2. I haven't seen the rule of this competition, e.g. Are we allowed to use pre-trained models or external data?\n",
    "1133631": "Thanks for the host and for the amazing dataset. \n\nKnowing some great kagglers are among organizers will surely help for quick and accurate feedback :)",
    "1233551": "@nguyenbadung @andy2709 @nguyenquyha \nThanks for hosting this interesting competition.\nI am wondering if semi-supervised learning methods such as pseudo-labeling is allowed in this competition? \nI believe it should be allowed, but I still want a confirmation from host. ",
    "1134118": "This is really an interesting problem and thanks for hosting this competition. ",
    "1563614": "Hello dear\nWhy MRI images are not used instead of X-rays to diagnose lung disease? \nAren't MRI images more accurate for this?",
    "1462434": "Hi ! I am new to Kaggle and trying to explore this competition and i am little confused about the values mentioned under Score within Leaderboard.\nSo, are these MAP (Mean average precision) values of each model? or is this something else?\nthank you.",
    "1457267": "i agree with you, kind sir\n",
    "1256347": "@nguyenquyha , @andy2709 , @nguyenbadung ,\nnow, that the end of the competition is near, i would like to ask you something technical.\nI found some attenuation patterns on many of the images: namely hyperlucent hemithoraces, mainly on the left side like here :\n![att](https://i.postimg.cc/RhDdzZJ8/attenuation.jpg)\nThese are not oblique chest radiographs - that could explain this phenomenon -, we can furthermore exclude the other obvious causes (absence of a breast after mastectomy, absence of a pectoralis muscle, unilateral or asymmetric bullous disease/emphysema, or air trapping, or an endobronchial foreign body). It made me really hard to evaluate the lung parenchyma!\n\nIs that a general technical issue? I must admit i do not encounter such images in our daily routine.\n\nThank you in advance.\n\n\n",
    "1192268": "Hi Ha,\nI'm wondering how can I access the age of the cases?\nThank you.",
    "1145500": "@nguyenquyha Thanks for organizing interesting challenge!\nI may miss to read some description, but may I ask how train/test set were split? ",
    "1732595": "Thanks for clarification.",
    "1143144": "Thanks for organizing this competition! "
  }
}