{
  "id": 211035,
  "title": "5 Radiologists' \"Consensus\"",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/211035",
  "author_name": "quillio",
  "post_date": "2021-01-13T11:27:55.454000",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>From the paper describing the dataset:</p>\n<blockquote>\n  <p>In this work, we describe a dataset of more than 100,000 chest X-ray scans that were retrospectively collected from two major hospitals in Vietnam. Out of this raw data, we release 18,000 images that were manually annotated by a total of 17 experienced radiologists with 22 local labels of rectangles surrounding abnormalities and 6 global labels of suspected diseases. The released dataset is divided into a training set of 15,000 and a test set of 3,000. Each scan in the training set was independently labeled by 3 radiologists, while <strong>each scan in the test set was labeled by the consensus of 5 radiologists</strong>. We designed and built a labeling platform for DICOM images to facilitate these annotation procedures. All images are made publicly available in DICOM format in company with the labels of the training set. The labels of the test set are hidden at the time of writing this paper as they will be used for benchmarking machine learning algorithms on an open platform.</p>\n</blockquote>\n<p>What do you think \"consensus\" means in this instance?</p>\n<p>This may matter for scoring in that the \"flavor\" of consensus may decrease the number of labels per image or may decrease the size of bounding boxes (in the test set relative to the training set).</p>\n<p>This would then bias us to prefer fewer labels and smaller bounding boxes, correct?</p>",
  "messages": [
    {
      "id": 1151504,
      "postDate": "2021-01-13T11:27:55.453Z",
      "content": "<p>From the paper describing the dataset:</p>\n<blockquote>\n  <p>In this work, we describe a dataset of more than 100,000 chest X-ray scans that were retrospectively collected from two major hospitals in Vietnam. Out of this raw data, we release 18,000 images that were manually annotated by a total of 17 experienced radiologists with 22 local labels of rectangles surrounding abnormalities and 6 global labels of suspected diseases. The released dataset is divided into a training set of 15,000 and a test set of 3,000. Each scan in the training set was independently labeled by 3 radiologists, while <strong>each scan in the test set was labeled by the consensus of 5 radiologists</strong>. We designed and built a labeling platform for DICOM images to facilitate these annotation procedures. All images are made publicly available in DICOM format in company with the labels of the training set. The labels of the test set are hidden at the time of writing this paper as they will be used for benchmarking machine learning algorithms on an open platform.</p>\n</blockquote>\n<p>What do you think \"consensus\" means in this instance?</p>\n<p>This may matter for scoring in that the \"flavor\" of consensus may decrease the number of labels per image or may decrease the size of bounding boxes (in the test set relative to the training set).</p>\n<p>This would then bias us to prefer fewer labels and smaller bounding boxes, correct?</p>",
      "rawMarkdown": "From the paper describing the dataset:\n> In this work, we describe a dataset of more than 100,000 chest X-ray scans that were retrospectively collected from two major hospitals in Vietnam. Out of this raw data, we release 18,000 images that were manually annotated by a total of 17 experienced radiologists with 22 local labels of rectangles surrounding abnormalities and 6 global labels of suspected diseases. The released dataset is divided into a training set of 15,000 and a test set of 3,000. Each scan in the training set was independently labeled by 3 radiologists, while **each scan in the test set was labeled by the consensus of 5 radiologists**. We designed and built a labeling platform for DICOM images to facilitate these annotation procedures. All images are made publicly available in DICOM format in company with the labels of the training set. The labels of the test set are hidden at the time of writing this paper as they will be used for benchmarking machine learning algorithms on an open platform.\n\nWhat do you think \"consensus\" means in this instance?\n\nThis may matter for scoring in that the \"flavor\" of consensus may decrease the number of labels per image or may decrease the size of bounding boxes (in the test set relative to the training set).\n\nThis would then bias us to prefer fewer labels and smaller bounding boxes, correct?",
      "votes": 5
    },
    {
      "id": 1152044,
      "postDate": "2021-01-13T18:17:36.483Z",
      "content": "<p>This quote from their paper may answer your question:</p>\n<blockquote>\n  <p>For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other<br>\n  radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.</p>\n</blockquote>\n<p>In my experience with this type of process (been involved in similar adjudication arrangements for clinical trials, but not with such a pure radiology focus), I would imagine that they have a discussion where they present arguments for what they think. E.g. if one of them missed something and the others point it out, I'd expect that to end up being labelled/a box to be expanded accordingly. On the other hand, there can indeed be a tendency towards not reaching a consensus on uncertain cases (=presumably would not become part of the ground truth). On a whole, I would think it can go either way (perhaps with a slight tend towards fewer/smaller boxes?!? I'm certainly not sure on that prediction…).</p>",
      "rawMarkdown": "This quote from their paper may answer your question:\n\n> For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other\nradiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\n\nIn my experience with this type of process (been involved in similar adjudication arrangements for clinical trials, but not with such a pure radiology focus), I would imagine that they have a discussion where they present arguments for what they think. E.g. if one of them missed something and the others point it out, I'd expect that to end up being labelled/a box to be expanded accordingly. On the other hand, there can indeed be a tendency towards not reaching a consensus on uncertain cases (=presumably would not become part of the ground truth). On a whole, I would think it can go either way (perhaps with a slight tend towards fewer/smaller boxes?!? I'm certainly not sure on that prediction...).",
      "votes": 2,
      "replies": [
        {
          "id": 1152693,
          "postDate": "2021-01-14T11:20:57.583Z",
          "content": "<p>My wondering comes in where we see a lot of bounding boxes overlap:<br>\nLet's say that the 3 radiologists in the training set put bounding boxes over the heart to signify \"aortic enlargement.\"  I think we all can assume that in the test, those bounding boxes will be combined into one ground truth box.</p>\n<p>But: what about other findings? Can there only be one \"cardiomegaly\"?  There definitely can be more than one \"Nodule/Mass.\"</p>\n<p>When and how to combine bounding boxes seems important.</p>",
          "rawMarkdown": "My wondering comes in where we see a lot of bounding boxes overlap:\nLet's say that the 3 radiologists in the training set put bounding boxes over the heart to signify \"aortic enlargement.\"  I think we all can assume that in the test, those bounding boxes will be combined into one ground truth box.\n\nBut: what about other findings? Can there only be one \"cardiomegaly\"?  There definitely can be more than one \"Nodule/Mass.\"\n\nWhen and how to combine bounding boxes seems important.",
          "votes": 2
        },
        {
          "id": 1152764,
          "postDate": "2021-01-14T12:37:55.360Z",
          "content": "<p>Yes, I'd have thought that certain things - when you look at their definition in the paper of some of the EDA notebooks that summarize this - should probably only have one bounding box, e.g. Aortic enlargement, Cardiomegaly. However, even for the some radiologist on the training data assign two boxes. I had already looked at this a bit before (see Section 5 of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/\" target=\"_blank\">this notebook</a>), but here's some more numbers)</p>\n<p>Maximum number of boxes:</p>\n<table>\n<thead>\n<tr>\n<th>class_name</th>\n<th>max boxes by one radiologist</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Nodule/Mass</td>\n<td>28</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>12</td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>11</td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>9</td>\n</tr>\n<tr>\n<td>Pleural thickening</td>\n<td>7</td>\n</tr>\n<tr>\n<td>Pulmonary fibrosis</td>\n<td>7</td>\n</tr>\n<tr>\n<td>Atelectasis</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Consolidation</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Infiltration</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Pneumothorax</td>\n<td>4</td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>3</td>\n</tr>\n<tr>\n<td>Pleural effusion</td>\n<td>2</td>\n</tr>\n<tr>\n<td>Aortic enlargement</td>\n<td>2</td>\n</tr>\n<tr>\n<td>Cardiomegaly</td>\n<td>2</td>\n</tr>\n<tr>\n<td>No finding</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>The mean number may be another indication. </p>\n<table>\n<thead>\n<tr>\n<th>class_name</th>\n<th>mean boxes by one radiologist (once at least one indicated)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No finding</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Aortic enlargement</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Cardiomegaly</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Atelectasis</td>\n<td>1.05</td>\n</tr>\n<tr>\n<td>Consolidation</td>\n<td>1.10</td>\n</tr>\n<tr>\n<td>Pneumothorax</td>\n<td>1.14</td>\n</tr>\n<tr>\n<td>Pleural effusion</td>\n<td>1.18</td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>1.22</td>\n</tr>\n<tr>\n<td>Infiltration</td>\n<td>1.31</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>1.36</td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>1.40</td>\n</tr>\n<tr>\n<td>Pulmonary fibrosis</td>\n<td>1.42</td>\n</tr>\n<tr>\n<td>Pleural thickening</td>\n<td>1.50</td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>1.64</td>\n</tr>\n<tr>\n<td>Nodule/Mass</td>\n<td>1.83</td>\n</tr>\n</tbody>\n</table>\n<p>I'd guess that the radiologists feel that all but cardiomegaly and aortic enlargement can occur more than once, while for those two classes, I'm wondering whether the few cases of those two having 2 bounding boxes are mistakes. </p>",
          "rawMarkdown": "Yes, I'd have thought that certain things - when you look at their definition in the paper of some of the EDA notebooks that summarize this - should probably only have one bounding box, e.g. Aortic enlargement, Cardiomegaly. However, even for the some radiologist on the training data assign two boxes. I had already looked at this a bit before (see Section 5 of [this notebook](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/)), but here's some more numbers)\n\nMaximum number of boxes:\n\n| class_name | max boxes by one radiologist |\n| --- | --- |\n| Nodule/Mass | 28 |\n| Other lesion | 12 |\n| Calcification | 11 |\n| Lung Opacity | 9 |\n| Pleural thickening | 7 |\n| Pulmonary fibrosis | 7 |\n| Atelectasis | 4 |\n| Consolidation | 4 |\n| Infiltration | 4 |\n| Pneumothorax | 4 |\n| ILD | 3 |\n| Pleural effusion | 2 |\n| Aortic enlargement | 2 |\n| Cardiomegaly | 2 |\n| No finding | 1 |\n\nThe mean number may be another indication. \n| class_name | mean boxes by one radiologist (once at least one indicated) |\t\n| --- | --- |\n| No finding | 1.00 |\n| Aortic enlargement | 1.00 |\n| Cardiomegaly | 1.00 |\n| Atelectasis | 1.05 |\n| Consolidation | 1.10 |\n| Pneumothorax | 1.14 |\n| Pleural effusion | 1.18 |\n| Lung Opacity | 1.22 |\n| Infiltration | 1.31 |\n| Other lesion | 1.36 |\n| Calcification | 1.40 |\n| Pulmonary fibrosis | 1.42 |\n| Pleural thickening | 1.50 |\n| ILD | 1.64 |\n| Nodule/Mass | 1.83 |\n\nI'd guess that the radiologists feel that all but cardiomegaly and aortic enlargement can occur more than once, while for those two classes, I'm wondering whether the few cases of those two having 2 bounding boxes are mistakes. ",
          "votes": 2
        },
        {
          "id": 1152866,
          "postDate": "2021-01-14T14:12:42.813Z",
          "content": "<p>Yes, hard to believe you can have two Aortic enlargements unless you have two hearts? 😂</p>",
          "rawMarkdown": "Yes, hard to believe you can have two Aortic enlargements unless you have two hearts? 😂",
          "votes": 1
        },
        {
          "id": 1152875,
          "postDate": "2021-01-14T14:23:04.787Z",
          "content": "<p>I suppose your aorta (largest artery) and similarly your heart for cariomegaly could be enlarged in two places or something, and the radiologist highlighted both??</p>",
          "rawMarkdown": "I suppose your aorta (largest artery) and similarly your heart for cariomegaly could be enlarged in two places or something, and the radiologist highlighted both??",
          "votes": 1
        },
        {
          "id": 1154750,
          "postDate": "2021-01-15T21:22:22.090Z",
          "content": "<p>Hello, I would like to make a small precision. Cardiomegaly and aortic enlargments have precise definitions and correspond to precise regions. For example, cardiomegaly corresponds to the increase in the size of the heart representing more than half of the thorax and is related to the whole heart. It is the same for aortic enlargements. I therefore don't understand the 2 boxes for this 2 items.<br>\nThe other items are ok. 4 localisations of pneumothorax are possible but it would be really really rare…<br>\nI will look for the corresponding images to understand and make another post !</p>",
          "rawMarkdown": "Hello, I would like to make a small precision. Cardiomegaly and aortic enlargments have precise definitions and correspond to precise regions. For example, cardiomegaly corresponds to the increase in the size of the heart representing more than half of the thorax and is related to the whole heart. It is the same for aortic enlargements. I therefore don't understand the 2 boxes for this 2 items.\nThe other items are ok. 4 localisations of pneumothorax are possible but it would be really really rare...\nI will look for the corresponding images to understand and make another post !",
          "votes": 1
        },
        {
          "id": 1154808,
          "postDate": "2021-01-15T23:28:36.187Z",
          "content": "<p>Here's some of the examples (note that <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">I remapped</a> the shortest dimension to 600 pixels to make things easier to manage - so it will not match up with <code>train.csv</code>).</p>\n<p>First, cardiomegaly examples (the bounding boxes are red, the ones for any other diagnosis blue)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fed65b86112a051019d9aa67a2c357576%2Fexample1.png?generation=1610752034001436&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F5413a6920caac6cc4392faaa6d9ca25d%2Fexample2.png?generation=1610752116276781&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fe72165359f091dbc8fc9e1c014401e8c%2Fexample3.png?generation=1610752131151438&amp;alt=media\" alt=\"\"><br>\nSecondly, aortic enlargement (yellow bounding boxes, the ones for any other diagnoses in blue):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Faf0e6b2aa5d6017586f11c91b6127eaa%2Fexample4.png?generation=1610752148102844&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F6b3dc47d3710e536a4e67ffb3d21bab3%2Fexample5.png?generation=1610752167088810&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffacaa4aa05ea1a500d689841f95b7424%2Fexample6.png?generation=1610752183709115&amp;alt=media\" alt=\"\"></p>\n<p>I'm no radiologist and I'm just comparing these to other examples, but am I right in guessing that these really are cases where there is both aortic enlargment and cardiomegaly present in the same picture, but the radiologist has accidentally annotated both as one of the two classes? That kind of makes sense, because those two findings (<code>class_id</code>s 0 and 3 in the plot below) coincide a lot:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffa2bc349c3cd5291bfa61bb6762cee56%2FScreenshot%20from%202021-01-16%2000-26-55.png?generation=1610753261467945&amp;alt=media\" alt=\"\"></p>\n<p>You can see the full list of examples and plots for all of those in the \"Possible data issues\" Section of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">this notebook (that also contains the correlation plot above and some other EDA)</a>, once it's finished running.</p>",
          "rawMarkdown": "Here's some of the examples (note that [I remapped](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray) the shortest dimension to 600 pixels to make things easier to manage - so it will not match up with `train.csv`).\n\nFirst, cardiomegaly examples (the bounding boxes are red, the ones for any other diagnosis blue)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fed65b86112a051019d9aa67a2c357576%2Fexample1.png?generation=1610752034001436&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F5413a6920caac6cc4392faaa6d9ca25d%2Fexample2.png?generation=1610752116276781&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fe72165359f091dbc8fc9e1c014401e8c%2Fexample3.png?generation=1610752131151438&alt=media)\nSecondly, aortic enlargement (yellow bounding boxes, the ones for any other diagnoses in blue):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Faf0e6b2aa5d6017586f11c91b6127eaa%2Fexample4.png?generation=1610752148102844&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F6b3dc47d3710e536a4e67ffb3d21bab3%2Fexample5.png?generation=1610752167088810&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffacaa4aa05ea1a500d689841f95b7424%2Fexample6.png?generation=1610752183709115&alt=media)\n\nI'm no radiologist and I'm just comparing these to other examples, but am I right in guessing that these really are cases where there is both aortic enlargment and cardiomegaly present in the same picture, but the radiologist has accidentally annotated both as one of the two classes? That kind of makes sense, because those two findings (`class_id`s 0 and 3 in the plot below) coincide a lot:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffa2bc349c3cd5291bfa61bb6762cee56%2FScreenshot%20from%202021-01-16%2000-26-55.png?generation=1610753261467945&alt=media)\n\nYou can see the full list of examples and plots for all of those in the \"Possible data issues\" Section of [this notebook (that also contains the correlation plot above and some other EDA)](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray), once it's finished running.",
          "votes": 3
        },
        {
          "id": 1154840,
          "postDate": "2021-01-16T01:39:44.730Z",
          "content": "<p>Seems like one can have an enlarged aorta and an enlarged heart at the same time. It makes sense that those might coincide.</p>",
          "rawMarkdown": "Seems like one can have an enlarged aorta and an enlarged heart at the same time. It makes sense that those might coincide.",
          "votes": 1
        },
        {
          "id": 1155063,
          "postDate": "2021-01-16T07:10:47.923Z",
          "content": "<p>Ho thanks <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ! Yes I think you are right, the radiologist has made a mistake …</p>",
          "rawMarkdown": "Ho thanks @bjoernholzhauer ! Yes I think you are right, the radiologist has made a mistake ...",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1152044,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-01-13T18:17:36.483000",
      "content": "<p>This quote from their paper may answer your question:</p>\n<blockquote>\n  <p>For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other<br>\n  radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.</p>\n</blockquote>\n<p>In my experience with this type of process (been involved in similar adjudication arrangements for clinical trials, but not with such a pure radiology focus), I would imagine that they have a discussion where they present arguments for what they think. E.g. if one of them missed something and the others point it out, I'd expect that to end up being labelled/a box to be expanded accordingly. On the other hand, there can indeed be a tendency towards not reaching a consensus on uncertain cases (=presumably would not become part of the ground truth). On a whole, I would think it can go either way (perhaps with a slight tend towards fewer/smaller boxes?!? I'm certainly not sure on that prediction…).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1152693,
          "author_name": "quillio",
          "author_url": "",
          "post_date": "2021-01-14T11:20:57.583000",
          "content": "<p>My wondering comes in where we see a lot of bounding boxes overlap:<br>\nLet's say that the 3 radiologists in the training set put bounding boxes over the heart to signify \"aortic enlargement.\"  I think we all can assume that in the test, those bounding boxes will be combined into one ground truth box.</p>\n<p>But: what about other findings? Can there only be one \"cardiomegaly\"?  There definitely can be more than one \"Nodule/Mass.\"</p>\n<p>When and how to combine bounding boxes seems important.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1152764,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-14T12:37:55.360000",
          "content": "<p>Yes, I'd have thought that certain things - when you look at their definition in the paper of some of the EDA notebooks that summarize this - should probably only have one bounding box, e.g. Aortic enlargement, Cardiomegaly. However, even for the some radiologist on the training data assign two boxes. I had already looked at this a bit before (see Section 5 of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/\" target=\"_blank\">this notebook</a>), but here's some more numbers)</p>\n<p>Maximum number of boxes:</p>\n<table>\n<thead>\n<tr>\n<th>class_name</th>\n<th>max boxes by one radiologist</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Nodule/Mass</td>\n<td>28</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>12</td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>11</td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>9</td>\n</tr>\n<tr>\n<td>Pleural thickening</td>\n<td>7</td>\n</tr>\n<tr>\n<td>Pulmonary fibrosis</td>\n<td>7</td>\n</tr>\n<tr>\n<td>Atelectasis</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Consolidation</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Infiltration</td>\n<td>4</td>\n</tr>\n<tr>\n<td>Pneumothorax</td>\n<td>4</td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>3</td>\n</tr>\n<tr>\n<td>Pleural effusion</td>\n<td>2</td>\n</tr>\n<tr>\n<td>Aortic enlargement</td>\n<td>2</td>\n</tr>\n<tr>\n<td>Cardiomegaly</td>\n<td>2</td>\n</tr>\n<tr>\n<td>No finding</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>The mean number may be another indication. </p>\n<table>\n<thead>\n<tr>\n<th>class_name</th>\n<th>mean boxes by one radiologist (once at least one indicated)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>No finding</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Aortic enlargement</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Cardiomegaly</td>\n<td>1.00</td>\n</tr>\n<tr>\n<td>Atelectasis</td>\n<td>1.05</td>\n</tr>\n<tr>\n<td>Consolidation</td>\n<td>1.10</td>\n</tr>\n<tr>\n<td>Pneumothorax</td>\n<td>1.14</td>\n</tr>\n<tr>\n<td>Pleural effusion</td>\n<td>1.18</td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>1.22</td>\n</tr>\n<tr>\n<td>Infiltration</td>\n<td>1.31</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>1.36</td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>1.40</td>\n</tr>\n<tr>\n<td>Pulmonary fibrosis</td>\n<td>1.42</td>\n</tr>\n<tr>\n<td>Pleural thickening</td>\n<td>1.50</td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>1.64</td>\n</tr>\n<tr>\n<td>Nodule/Mass</td>\n<td>1.83</td>\n</tr>\n</tbody>\n</table>\n<p>I'd guess that the radiologists feel that all but cardiomegaly and aortic enlargement can occur more than once, while for those two classes, I'm wondering whether the few cases of those two having 2 bounding boxes are mistakes. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1152866,
          "author_name": "quillio",
          "author_url": "",
          "post_date": "2021-01-14T14:12:42.813000",
          "content": "<p>Yes, hard to believe you can have two Aortic enlargements unless you have two hearts? 😂</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1152875,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-14T14:23:04.787000",
          "content": "<p>I suppose your aorta (largest artery) and similarly your heart for cariomegaly could be enlarged in two places or something, and the radiologist highlighted both??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1154750,
          "author_name": "Tanguy Perennec",
          "author_url": "",
          "post_date": "2021-01-15T21:22:22.090000",
          "content": "<p>Hello, I would like to make a small precision. Cardiomegaly and aortic enlargments have precise definitions and correspond to precise regions. For example, cardiomegaly corresponds to the increase in the size of the heart representing more than half of the thorax and is related to the whole heart. It is the same for aortic enlargements. I therefore don't understand the 2 boxes for this 2 items.<br>\nThe other items are ok. 4 localisations of pneumothorax are possible but it would be really really rare…<br>\nI will look for the corresponding images to understand and make another post !</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1154808,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-15T23:28:36.187000",
          "content": "<p>Here's some of the examples (note that <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">I remapped</a> the shortest dimension to 600 pixels to make things easier to manage - so it will not match up with <code>train.csv</code>).</p>\n<p>First, cardiomegaly examples (the bounding boxes are red, the ones for any other diagnosis blue)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fed65b86112a051019d9aa67a2c357576%2Fexample1.png?generation=1610752034001436&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F5413a6920caac6cc4392faaa6d9ca25d%2Fexample2.png?generation=1610752116276781&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fe72165359f091dbc8fc9e1c014401e8c%2Fexample3.png?generation=1610752131151438&amp;alt=media\" alt=\"\"><br>\nSecondly, aortic enlargement (yellow bounding boxes, the ones for any other diagnoses in blue):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Faf0e6b2aa5d6017586f11c91b6127eaa%2Fexample4.png?generation=1610752148102844&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F6b3dc47d3710e536a4e67ffb3d21bab3%2Fexample5.png?generation=1610752167088810&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffacaa4aa05ea1a500d689841f95b7424%2Fexample6.png?generation=1610752183709115&amp;alt=media\" alt=\"\"></p>\n<p>I'm no radiologist and I'm just comparing these to other examples, but am I right in guessing that these really are cases where there is both aortic enlargment and cardiomegaly present in the same picture, but the radiologist has accidentally annotated both as one of the two classes? That kind of makes sense, because those two findings (<code>class_id</code>s 0 and 3 in the plot below) coincide a lot:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Ffa2bc349c3cd5291bfa61bb6762cee56%2FScreenshot%20from%202021-01-16%2000-26-55.png?generation=1610753261467945&amp;alt=media\" alt=\"\"></p>\n<p>You can see the full list of examples and plots for all of those in the \"Possible data issues\" Section of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">this notebook (that also contains the correlation plot above and some other EDA)</a>, once it's finished running.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1154840,
          "author_name": "quillio",
          "author_url": "",
          "post_date": "2021-01-16T01:39:44.730000",
          "content": "<p>Seems like one can have an enlarged aorta and an enlarged heart at the same time. It makes sense that those might coincide.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1155063,
          "author_name": "Tanguy Perennec",
          "author_url": "",
          "post_date": "2021-01-16T07:10:47.923000",
          "content": "<p>Ho thanks <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ! Yes I think you are right, the radiologist has made a mistake …</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1151504": "From the paper describing the dataset:\n> In this work, we describe a dataset of more than 100,000 chest X-ray scans that were retrospectively collected from two major hospitals in Vietnam. Out of this raw data, we release 18,000 images that were manually annotated by a total of 17 experienced radiologists with 22 local labels of rectangles surrounding abnormalities and 6 global labels of suspected diseases. The released dataset is divided into a training set of 15,000 and a test set of 3,000. Each scan in the training set was independently labeled by 3 radiologists, while **each scan in the test set was labeled by the consensus of 5 radiologists**. We designed and built a labeling platform for DICOM images to facilitate these annotation procedures. All images are made publicly available in DICOM format in company with the labels of the training set. The labels of the test set are hidden at the time of writing this paper as they will be used for benchmarking machine learning algorithms on an open platform.\n\nWhat do you think \"consensus\" means in this instance?\n\nThis may matter for scoring in that the \"flavor\" of consensus may decrease the number of labels per image or may decrease the size of bounding boxes (in the test set relative to the training set).\n\nThis would then bias us to prefer fewer labels and smaller bounding boxes, correct?",
    "1152044": "This quote from their paper may answer your question:\n\n> For the test set, 5 radiologists involved into a two-stage labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other\nradiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\n\nIn my experience with this type of process (been involved in similar adjudication arrangements for clinical trials, but not with such a pure radiology focus), I would imagine that they have a discussion where they present arguments for what they think. E.g. if one of them missed something and the others point it out, I'd expect that to end up being labelled/a box to be expanded accordingly. On the other hand, there can indeed be a tendency towards not reaching a consensus on uncertain cases (=presumably would not become part of the ground truth). On a whole, I would think it can go either way (perhaps with a slight tend towards fewer/smaller boxes?!? I'm certainly not sure on that prediction...)."
  }
}