{
  "id": 212859,
  "title": "Mislabeled bounding boxes & other data issues (may need a radiologist here...)",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/212859",
  "author_name": "Björn",
  "post_date": "2021-01-20T13:53:47.702000",
  "votes": 14,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Based on some of the <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211035\" target=\"_blank\">discussions</a>, where we uncovered some mislabelled bounding boxes, that suggested that at least some 'aortic enlargement' and 'cardiomegaly' labels for bounding boxes are wrong, I realized that there's definitely some mislabeled bounding boxes.</p>\n<p>Should we care? I think so, if the training data is not generated from the same process as the test data (2 additional radiologists re-conciliate issues for the test data) and a lot of the differences will be mistakes in the training data that would likely be fixed, if the image were in the test data. Thus, it makes sense to try and fix the issues we can fix on the training data (that should improve the link from training and CV performance to LB).</p>\n<p>Thus, I tried one approach similar to how we uncovered issues in the forum discussions by using LightGBM on the <code>train.csv</code> and some additional <code>.dicom</code> meta-data. <strong>The notebook implementing this model and producing the images below is <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes/\" target=\"_blank\">here</a>.</strong></p>\n<p>Here are some examples of issues that I found. I'm not a radiologist, so take my human assessment of these with a grain of salt (expert input very, very welcome!). Firstly, the approach again picked up some of those  'aortic enlargement' vs. 'cardiomegaly' issues:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F2493608399172fec8083d7a821336ffc%2F__results___41_3.png?generation=1611149547487964&amp;alt=media\" alt=\"\"><br>\nI'm also pretty sure that cardiomegaly should not be in this location, so presumably the model is right that this could be aortic enlargement?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fc237c3c34fdac9358d347ce07d4883cb%2F__results___41_9.png?generation=1611149947019757&amp;alt=media\" alt=\"\"><br>\nHere's a case of aortic enlargement labelled in a weird position (the one on the left):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fa9738147a42ecb7d6e7a54c5a5b82ddf%2F__results___41_5.png?generation=1611149609228968&amp;alt=media\" alt=\"\"><br>\nHere's another example of a weird position for a class and very small bounding box by the standards of the class (so the aortic enlargement suggestion of the model does make sense to me):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F44dc74f173eacdeff9501f741a713b7a%2F__results___41_10.png?generation=1611150122673016&amp;alt=media\" alt=\"\"><br>\nBut sometimes, I just don't know what to make of it. E.g. is the model right that this is pleural thickening rather than pneumothorax (<a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#5.-How-big-do-bounding-boxes-tend-to-be-for-different-classes?-How-many-are-there?\" target=\"_blank\">EDA suggests</a> that pneumothorax tends to be very large bounding boxes and pleural thickening tends to be in <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?\" target=\"_blank\">this particular region</a> of x-rays with smaller bounding boxes, but I cannot really judge this well)? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fee3d330f108eb32b749b591b964f33de%2F__results___41_8.png?generation=1611149721332674&amp;alt=media\" alt=\"\"><br>\nI've identified at least one case, where I think my model is wrong (in the sense that it looks like the radiologist had already labelled an aortic enlargement and clearly did intend to label the overlapping - but not identical - bounding box as Nodule/Mass) - perhaps I could avoid this (presumably) wrong prediction by adding a feature counting other labels assigned by the same radiologist to other bounding boxes (<strong>Update:</strong> Adding the extra feature that I mentioned in version 17 of the notebook seems to have stopped the model from capturing the image I thought it should not identify):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fec64766573aa794fee39c841d7a0caa9%2F__results___41_46.png?generation=1611150416122729&amp;alt=media\" alt=\"\"><br>\nThe second (higher) box here seems wrong, but who knows whether the model is right that this is ILD:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F1f493c5d1c3d677038122836e748e839%2F__results___41_52.png?generation=1611150671637265&amp;alt=media\" alt=\"\"></p>\n<p>There's plenty more examples in the notebook, but a lot of the cases seem to involve aortic enlargement, cardiomegaly, pleural effusion/thickening and occasionally a few other classes.</p>\n<p>In any case, it looks like it might be worthwhile to try to put together a cleaned up dataset, but it's challenging as a non-radiologist.</p>",
  "messages": [
    {
      "id": 1161322,
      "postDate": "2021-01-20T13:53:47.703Z",
      "content": "<p>Based on some of the <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211035\" target=\"_blank\">discussions</a>, where we uncovered some mislabelled bounding boxes, that suggested that at least some 'aortic enlargement' and 'cardiomegaly' labels for bounding boxes are wrong, I realized that there's definitely some mislabeled bounding boxes.</p>\n<p>Should we care? I think so, if the training data is not generated from the same process as the test data (2 additional radiologists re-conciliate issues for the test data) and a lot of the differences will be mistakes in the training data that would likely be fixed, if the image were in the test data. Thus, it makes sense to try and fix the issues we can fix on the training data (that should improve the link from training and CV performance to LB).</p>\n<p>Thus, I tried one approach similar to how we uncovered issues in the forum discussions by using LightGBM on the <code>train.csv</code> and some additional <code>.dicom</code> meta-data. <strong>The notebook implementing this model and producing the images below is <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes/\" target=\"_blank\">here</a>.</strong></p>\n<p>Here are some examples of issues that I found. I'm not a radiologist, so take my human assessment of these with a grain of salt (expert input very, very welcome!). Firstly, the approach again picked up some of those  'aortic enlargement' vs. 'cardiomegaly' issues:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F2493608399172fec8083d7a821336ffc%2F__results___41_3.png?generation=1611149547487964&amp;alt=media\" alt=\"\"><br>\nI'm also pretty sure that cardiomegaly should not be in this location, so presumably the model is right that this could be aortic enlargement?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fc237c3c34fdac9358d347ce07d4883cb%2F__results___41_9.png?generation=1611149947019757&amp;alt=media\" alt=\"\"><br>\nHere's a case of aortic enlargement labelled in a weird position (the one on the left):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fa9738147a42ecb7d6e7a54c5a5b82ddf%2F__results___41_5.png?generation=1611149609228968&amp;alt=media\" alt=\"\"><br>\nHere's another example of a weird position for a class and very small bounding box by the standards of the class (so the aortic enlargement suggestion of the model does make sense to me):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F44dc74f173eacdeff9501f741a713b7a%2F__results___41_10.png?generation=1611150122673016&amp;alt=media\" alt=\"\"><br>\nBut sometimes, I just don't know what to make of it. E.g. is the model right that this is pleural thickening rather than pneumothorax (<a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#5.-How-big-do-bounding-boxes-tend-to-be-for-different-classes?-How-many-are-there?\" target=\"_blank\">EDA suggests</a> that pneumothorax tends to be very large bounding boxes and pleural thickening tends to be in <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?\" target=\"_blank\">this particular region</a> of x-rays with smaller bounding boxes, but I cannot really judge this well)? <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fee3d330f108eb32b749b591b964f33de%2F__results___41_8.png?generation=1611149721332674&amp;alt=media\" alt=\"\"><br>\nI've identified at least one case, where I think my model is wrong (in the sense that it looks like the radiologist had already labelled an aortic enlargement and clearly did intend to label the overlapping - but not identical - bounding box as Nodule/Mass) - perhaps I could avoid this (presumably) wrong prediction by adding a feature counting other labels assigned by the same radiologist to other bounding boxes (<strong>Update:</strong> Adding the extra feature that I mentioned in version 17 of the notebook seems to have stopped the model from capturing the image I thought it should not identify):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fec64766573aa794fee39c841d7a0caa9%2F__results___41_46.png?generation=1611150416122729&amp;alt=media\" alt=\"\"><br>\nThe second (higher) box here seems wrong, but who knows whether the model is right that this is ILD:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F1f493c5d1c3d677038122836e748e839%2F__results___41_52.png?generation=1611150671637265&amp;alt=media\" alt=\"\"></p>\n<p>There's plenty more examples in the notebook, but a lot of the cases seem to involve aortic enlargement, cardiomegaly, pleural effusion/thickening and occasionally a few other classes.</p>\n<p>In any case, it looks like it might be worthwhile to try to put together a cleaned up dataset, but it's challenging as a non-radiologist.</p>",
      "rawMarkdown": "Based on some of the [discussions](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211035), where we uncovered some mislabelled bounding boxes, that suggested that at least some 'aortic enlargement' and 'cardiomegaly' labels for bounding boxes are wrong, I realized that there's definitely some mislabeled bounding boxes.\n\nShould we care? I think so, if the training data is not generated from the same process as the test data (2 additional radiologists re-conciliate issues for the test data) and a lot of the differences will be mistakes in the training data that would likely be fixed, if the image were in the test data. Thus, it makes sense to try and fix the issues we can fix on the training data (that should improve the link from training and CV performance to LB).\n\nThus, I tried one approach similar to how we uncovered issues in the forum discussions by using LightGBM on the `train.csv` and some additional `.dicom` meta-data. **The notebook implementing this model and producing the images below is [here](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes/).**\n\nHere are some examples of issues that I found. I'm not a radiologist, so take my human assessment of these with a grain of salt (expert input very, very welcome!). Firstly, the approach again picked up some of those  'aortic enlargement' vs. 'cardiomegaly' issues:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F2493608399172fec8083d7a821336ffc%2F__results___41_3.png?generation=1611149547487964&alt=media)\nI'm also pretty sure that cardiomegaly should not be in this location, so presumably the model is right that this could be aortic enlargement?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fc237c3c34fdac9358d347ce07d4883cb%2F__results___41_9.png?generation=1611149947019757&alt=media)\nHere's a case of aortic enlargement labelled in a weird position (the one on the left):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fa9738147a42ecb7d6e7a54c5a5b82ddf%2F__results___41_5.png?generation=1611149609228968&alt=media)\nHere's another example of a weird position for a class and very small bounding box by the standards of the class (so the aortic enlargement suggestion of the model does make sense to me):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F44dc74f173eacdeff9501f741a713b7a%2F__results___41_10.png?generation=1611150122673016&alt=media)\nBut sometimes, I just don't know what to make of it. E.g. is the model right that this is pleural thickening rather than pneumothorax ([EDA suggests](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#5.-How-big-do-bounding-boxes-tend-to-be-for-different-classes?-How-many-are-there?) that pneumothorax tends to be very large bounding boxes and pleural thickening tends to be in [this particular region](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?) of x-rays with smaller bounding boxes, but I cannot really judge this well)? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fee3d330f108eb32b749b591b964f33de%2F__results___41_8.png?generation=1611149721332674&alt=media)\nI've identified at least one case, where I think my model is wrong (in the sense that it looks like the radiologist had already labelled an aortic enlargement and clearly did intend to label the overlapping - but not identical - bounding box as Nodule/Mass) - perhaps I could avoid this (presumably) wrong prediction by adding a feature counting other labels assigned by the same radiologist to other bounding boxes (**Update:** Adding the extra feature that I mentioned in version 17 of the notebook seems to have stopped the model from capturing the image I thought it should not identify):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fec64766573aa794fee39c841d7a0caa9%2F__results___41_46.png?generation=1611150416122729&alt=media)\nThe second (higher) box here seems wrong, but who knows whether the model is right that this is ILD:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F1f493c5d1c3d677038122836e748e839%2F__results___41_52.png?generation=1611150671637265&alt=media)\n\n\nThere's plenty more examples in the notebook, but a lot of the cases seem to involve aortic enlargement, cardiomegaly, pleural effusion/thickening and occasionally a few other classes.\n\nIn any case, it looks like it might be worthwhile to try to put together a cleaned up dataset, but it's challenging as a non-radiologist.",
      "votes": 14
    },
    {
      "id": 1211810,
      "postDate": "2021-02-20T15:26:02.103Z",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>,<br>\nradiologist here. It is also challenging for a radiologist =) </p>\n<p>To clean cardiomegaly and aortic enlargement is pretty straigtforward and easy job, but there are some classes that would make pretty huge inter-observer variance:  ILD vs pulmonary fibrosis… or consolidation / infiltration / opacity. </p>",
      "rawMarkdown": "@bjoernholzhauer,\nradiologist here. It is also challenging for a radiologist =) \n\nTo clean cardiomegaly and aortic enlargement is pretty straigtforward and easy job, but there are some classes that would make pretty huge inter-observer variance:  ILD vs pulmonary fibrosis... or consolidation / infiltration / opacity. \n",
      "votes": 5
    },
    {
      "id": 1162125,
      "postDate": "2021-01-21T02:23:01.440Z",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> nice finding! I was just about to investigate this too, but without using any machine learning. When you look at bounding box centers for a class of findings - say Aortic Enlargement - and then check for bounding box centers that are 3 or more standard deviations from the mean, you find the potential errors pretty quickly. My <a href=\"https://www.kaggle.com/craigmthomas/localization-of-findings\" target=\"_blank\">quick EDA</a> does this. It shows the density heatmap next to the bounding box centers. On the bounding box center plot it measures how many standard deviations from the mean each bounding box center appears. This is for Aortic Enlargement:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2F70b393afd95cd06ce5d2f1efea60c87b%2Faortic_enlargment_bbox_centers.png?generation=1611195204811621&amp;alt=media\" alt=\"\"></p>\n<p>The red dots trailing down near the center are more than 4 standard deviations from the mean. I took a quick look at the actual bounding box for some of them, and my initial thoughts were that they must have been Cardiomegaly that were accidentally mislabeled as Aortic Enlargement (but I'm not a radiologist, so I can't say that definitively). Your machine learning approach for detection is pretty neat. Thanks for sharing!</p>",
      "rawMarkdown": "@bjoernholzhauer nice finding! I was just about to investigate this too, but without using any machine learning. When you look at bounding box centers for a class of findings - say Aortic Enlargement - and then check for bounding box centers that are 3 or more standard deviations from the mean, you find the potential errors pretty quickly. My [quick EDA](https://www.kaggle.com/craigmthomas/localization-of-findings) does this. It shows the density heatmap next to the bounding box centers. On the bounding box center plot it measures how many standard deviations from the mean each bounding box center appears. This is for Aortic Enlargement:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2F70b393afd95cd06ce5d2f1efea60c87b%2Faortic_enlargment_bbox_centers.png?generation=1611195204811621&alt=media)\n\nThe red dots trailing down near the center are more than 4 standard deviations from the mean. I took a quick look at the actual bounding box for some of them, and my initial thoughts were that they must have been Cardiomegaly that were accidentally mislabeled as Aortic Enlargement (but I'm not a radiologist, so I can't say that definitively). Your machine learning approach for detection is pretty neat. Thanks for sharing!",
      "votes": 6,
      "replies": [
        {
          "id": 1211813,
          "postDate": "2021-02-20T15:29:35.243Z",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a>,</p>\n<p>nice visualisation there!<br>\nThe problem with the absolute position and it's deviation from the mean is that sometimes (not to rarely) the positioning of the patient i off-center, causing a shift of all organ projections on the image.</p>",
          "rawMarkdown": "@craigmthomas,\n\nnice visualisation there!\nThe problem with the absolute position and it's deviation from the mean is that sometimes (not to rarely) the positioning of the patient i off-center, causing a shift of all organ projections on the image.",
          "votes": 4
        },
        {
          "id": 1211979,
          "postDate": "2021-02-20T18:05:32.120Z",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> thank you, your comment is very helpful! Being a radiologist, I'm sure you've definitely seen a lot of different x-ray alignments. In fact, perhaps you may be able to give us an idea about how many rotational degrees we could apply to our training set to achieve a realistic set of transforms to bolster our data (i.e. a rotation of 2 - 5 degrees may be appropriate, but 180 degrees is not likely to occur)? As well, maybe insights into the X and Y skews that may be appropriate?</p>\n<p>As to my visualization, I should have prefaced that finding with the fact that most of those &lt; 1% occurrences should be investigated manually to see if they are misclassifications (definitely have to be careful with applying any kind of automated cleaning). Most certainly, deviation from the mean may not be the best way to detect outliers in this context, but box centers occurring more than 4 standard deviations are worth investigating. With that said, I did that for Aortic Enlargment and found that some of the lower red outliers look very much like boxes for Cardiomegaly. There were also some higher up that looked like Pleural Thickening. </p>\n<p>As per your comment, I found the utility of doing this outside of Aortic Enlargement and Cardiomegaly is limited. Both Aortic Enlargement and Cardiomegaly tend to be very tightly grouped in the test data. Certainly when you get into Pleural Effusion and other classes, absolute location isn't a reliable indicator of mislabelled boxes, and random shifts in patient alignment have a much greater impact on what could potentially be an outlier. </p>\n<p>Thanks again for the comment, I really appreciate it!</p>",
          "rawMarkdown": "@sandorkonya thank you, your comment is very helpful! Being a radiologist, I'm sure you've definitely seen a lot of different x-ray alignments. In fact, perhaps you may be able to give us an idea about how many rotational degrees we could apply to our training set to achieve a realistic set of transforms to bolster our data (i.e. a rotation of 2 - 5 degrees may be appropriate, but 180 degrees is not likely to occur)? As well, maybe insights into the X and Y skews that may be appropriate?\n\nAs to my visualization, I should have prefaced that finding with the fact that most of those < 1% occurrences should be investigated manually to see if they are misclassifications (definitely have to be careful with applying any kind of automated cleaning). Most certainly, deviation from the mean may not be the best way to detect outliers in this context, but box centers occurring more than 4 standard deviations are worth investigating. With that said, I did that for Aortic Enlargment and found that some of the lower red outliers look very much like boxes for Cardiomegaly. There were also some higher up that looked like Pleural Thickening. \n\nAs per your comment, I found the utility of doing this outside of Aortic Enlargement and Cardiomegaly is limited. Both Aortic Enlargement and Cardiomegaly tend to be very tightly grouped in the test data. Certainly when you get into Pleural Effusion and other classes, absolute location isn't a reliable indicator of mislabelled boxes, and random shifts in patient alignment have a much greater impact on what could potentially be an outlier. \n\nThanks again for the comment, I really appreciate it!",
          "votes": 1
        },
        {
          "id": 1212695,
          "postDate": "2021-02-21T13:44:37.597Z",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a> ,<br>\nthe question of the augmentation is very interesting and not easily to answer. By the chest x-rays the \"rotation around the midpoint of the image\" occurs rather in only a few degrees range, since this is something what the assistant sees and controls during the acquisition! More problematic is the rotation of the patient of his own axis and the projection's center point ( imagine moving the radiation source in y axis as seen on the second image). Both of them changes the projection of all organs and the appearance of the pathologies and normal anatomy to.</p>\n<p><img src=\"https://i.postimg.cc/rpf5NW4m/xray4-1.jpg\" alt=\"chest\"> </p>\n<p>There are cases where the images are upside down or rotated to x*π/2 =)</p>\n<p>Nono, the method is VERY good, in fact i would do something similar to select outliers to look for for the first time cleaning. I know that there are many false labelings in the set. Maybe i find the time and annotate a random 5k aortic knob (from the non enlarged ones), so one will be able to train a \"aortic knob detection\", would be a good accompanying set to <a href=\"https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray\" target=\"_blank\">this one</a>.</p>\n<p>Cooccurence of cardiovascular diseases is a known fact, nice to see the theory in practice here =)</p>",
          "rawMarkdown": "@craigmthomas ,\nthe question of the augmentation is very interesting and not easily to answer. By the chest x-rays the \"rotation around the midpoint of the image\" occurs rather in only a few degrees range, since this is something what the assistant sees and controls during the acquisition! More problematic is the rotation of the patient of his own axis and the projection's center point ( imagine moving the radiation source in y axis as seen on the second image). Both of them changes the projection of all organs and the appearance of the pathologies and normal anatomy to.\n\n![chest](https://i.postimg.cc/rpf5NW4m/xray4-1.jpg) \n\nThere are cases where the images are upside down or rotated to x*π/2 =)\n\nNono, the method is VERY good, in fact i would do something similar to select outliers to look for for the first time cleaning. I know that there are many false labelings in the set. Maybe i find the time and annotate a random 5k aortic knob (from the non enlarged ones), so one will be able to train a \"aortic knob detection\", would be a good accompanying set to [this one](https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray).\n\nCooccurence of cardiovascular diseases is a known fact, nice to see the theory in practice here =)\n",
          "votes": 2
        },
        {
          "id": 1213313,
          "postDate": "2021-02-22T03:42:25.260Z",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> thank you so much for your comment! Your insight about the projection angle changing both normal and pathological appearance is very interesting. As well, the idea of an aortic knob detector is also quite interesting. Having the ability to create a waypoint like that on the image may help. I've been working on an ensemble solution, one where there is a classifier and bounding box regressor for each class of finding. Each classifier and regressor tries to focus on specific areas of the x-ray. The classifier portion works well for aortic enlargement and cardiomegaly because they tend to appear in roughly the same absolute area (the regressors are more of a challenge due to ground truth). Overall just focusing on these two classes has yielded a LB score of 0.08, however, identifying features to use purely for positioning purposes could potentially help out a lot - they could be used zoom in on the target areas more accurately. Thanks again for your comments, I really appreciate them!</p>",
          "rawMarkdown": "@sandorkonya thank you so much for your comment! Your insight about the projection angle changing both normal and pathological appearance is very interesting. As well, the idea of an aortic knob detector is also quite interesting. Having the ability to create a waypoint like that on the image may help. I've been working on an ensemble solution, one where there is a classifier and bounding box regressor for each class of finding. Each classifier and regressor tries to focus on specific areas of the x-ray. The classifier portion works well for aortic enlargement and cardiomegaly because they tend to appear in roughly the same absolute area (the regressors are more of a challenge due to ground truth). Overall just focusing on these two classes has yielded a LB score of 0.08, however, identifying features to use purely for positioning purposes could potentially help out a lot - they could be used zoom in on the target areas more accurately. Thanks again for your comments, I really appreciate them!",
          "votes": 2
        },
        {
          "id": 1217842,
          "postDate": "2021-02-25T11:09:33.047Z",
          "content": "<p>The idea with the regressor / classifier is good, but as I was browsing the labelled images, I found several classes to have pretty arbitrarily placed ROIs…<br>\nEspecially in case of ILD &amp;Fibrosis - these are the 2 where one can argue why they chose especially the position for that (see for example your 6th image).<br>\nIn case of these 2 labels it would be the best not to take the labels of the GT but the whole lung area (like in the second example), and create a classifier for the disease and then analyze the distribution of GT for the most likely position &amp; size for these bounding boxes. As you see on the second example, the classes ILD &amp; Fibrosis are not even separated… cooccur on the same bbox.</p>",
          "rawMarkdown": "The idea with the regressor / classifier is good, but as I was browsing the labelled images, I found several classes to have pretty arbitrarily placed ROIs...\nEspecially in case of ILD &Fibrosis - these are the 2 where one can argue why they chose especially the position for that (see for example your 6th image).\nIn case of these 2 labels it would be the best not to take the labels of the GT but the whole lung area (like in the second example), and create a classifier for the disease and then analyze the distribution of GT for the most likely position & size for these bounding boxes. As you see on the second example, the classes ILD & Fibrosis are not even separated... cooccur on the same bbox.",
          "votes": 2
        },
        {
          "id": 1220776,
          "postDate": "2021-02-28T10:55:31.787Z",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a>,</p>\n<p>FYI <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222626\" target=\"_blank\">this</a>.<br>\nExactly this is the way how the pathologies also \"change\" their appearance depending on the projection. Nevertheless this is a problem for mainly the longitudinal data analysis (in the current competition we have no duplicate of patients) but it is surely a factor that influences the inter-observer  error.</p>\n<p>I reviewed (…okay, that's to much…  i just scrolled through) the 15k training data and there are ~few dozens of images with an appearance of a portable/bedside image so in this competition this kind of bias play rather a smaller role!</p>",
          "rawMarkdown": "@craigmthomas,\n\nFYI [this](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222626).\nExactly this is the way how the pathologies also \"change\" their appearance depending on the projection. Nevertheless this is a problem for mainly the longitudinal data analysis (in the current competition we have no duplicate of patients) but it is surely a factor that influences the inter-observer  error.\n\nI reviewed (...okay, that's to much...  i just scrolled through) the 15k training data and there are ~few dozens of images with an appearance of a portable/bedside image so in this competition this kind of bias play rather a smaller role!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1161399,
      "postDate": "2021-01-20T14:49:36.913Z",
      "content": "<p>I found another annotation problem.<br>\nMultiple labels are assigned to exact same bbox, this can be found in 1775 images (mostly annotated by R8, R9, R10).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1111209%2F43e93cd1401dfd1df532fdd1f37473de%2Fss.jpg?generation=1611154145439806&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I found another annotation problem.\nMultiple labels are assigned to exact same bbox, this can be found in 1775 images (mostly annotated by R8, R9, R10).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1111209%2F43e93cd1401dfd1df532fdd1f37473de%2Fss.jpg?generation=1611154145439806&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 1161434,
          "postDate": "2021-01-20T15:10:48.727Z",
          "content": "<p>Interesting. I wonder whether that is actually something that makes sense or is a data issues… E.g. infiltration and pulmonary fibrosis co-occur a lot (correlation 0.38), so if the infiltration is what the radiologist thinks is what identifies the pulmonary fibrosis, then maybe this is sensible?!</p>",
          "rawMarkdown": "Interesting. I wonder whether that is actually something that makes sense or is a data issues... E.g. infiltration and pulmonary fibrosis co-occur a lot (correlation 0.38), so if the infiltration is what the radiologist thinks is what identifies the pulmonary fibrosis, then maybe this is sensible?!",
          "votes": 2
        },
        {
          "id": 1161537,
          "postDate": "2021-01-20T15:57:45.503Z",
          "content": "<p>Is it a problem? This is also known as multi-label classification.</p>",
          "rawMarkdown": "Is it a problem? This is also known as multi-label classification."
        },
        {
          "id": 1161564,
          "postDate": "2021-01-20T16:10:31.460Z",
          "content": "<p>Not a problem, if it's correct…</p>",
          "rawMarkdown": "Not a problem, if it's correct..."
        },
        {
          "id": 1161967,
          "postDate": "2021-01-20T21:24:15.193Z",
          "content": "<p>This should not be a problem with annotations as stated also from competition host in <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/209298\" target=\"_blank\">this thread</a></p>\n<pre><code>Some dense lesion area contains many labels, our radiologists created a box then assign many labels to that box.\n</code></pre>",
          "rawMarkdown": "This should not be a problem with annotations as stated also from competition host in [this thread](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/209298)\n\n```\nSome dense lesion area contains many labels, our radiologists created a box then assign many labels to that box.\n```",
          "votes": 3
        },
        {
          "id": 1162067,
          "postDate": "2021-01-21T00:04:39.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> Thank you ! I missed the thread.</p>",
          "rawMarkdown": "@lazcoder Thank you ! I missed the thread.",
          "votes": 1
        },
        {
          "id": 1162617,
          "postDate": "2021-01-21T08:36:58.240Z",
          "content": "<p>I've <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes#Bounding-boxes-with-more-than-one-label\" target=\"_blank\">summarized those cases</a> now and I guess I'm not sure whether you would call all of those \"dense lesion area\"(s)?  I'm not saying that these don't mostly make sense, but rather that there might be other patterns (there's only very few, were I'm semi-confident to say they are wrong).</p>\n<p>E.g. I'm not sure whether one would describe the most common combination of pleural effusion and thickening as lesions (perhaps?)? Something like <em>0: Aortic enlargement + 2: Calcification</em> also does not exactly fit the pattern they stated (or would you call calcification a \"lesion\"?). I guess the cases below in <strong>bold</strong> that are more \"something you see\" (quite possibly a dense lesion area) + \"assigned disease diagnosis\" (such as pulmonary fibrosis or ILD) arguably fit the pattern (just the dense lesion area gets assigned a disease diagnosis, too).</p>\n<p>The below are all the combinations that you have for at least 10 bounding boxes (there's obviously tons of rare combinations that happen for just 1), the number in brackets is the number of bounding boxes with this combination of labels:</p>\n<ul>\n<li>10: Pleural effusion + 11: Pleural thickening (1285)</li>\n<li><strong>6: Infiltration + 13: Pulmonary fibrosis (318)</strong></li>\n<li>6: Infiltration + 7: Lung Opacity (216)</li>\n<li>7: Lung Opacity + 8: Nodule/Mass (161)</li>\n<li>7: Lung Opacity + 9: Other lesion (158)</li>\n<li>4: Consolidation + 7: Lung Opacity (155)</li>\n<li><strong>2: Calcification + 13: Pulmonary fibrosis (153)</strong></li>\n<li><strong>7: Lung Opacity + 13: Pulmonary fibrosis (138)</strong></li>\n<li>8: Nodule/Mass + 9: Other lesion (136)</li>\n<li><strong>9: Other lesion + 13: Pulmonary fibrosis (106)</strong></li>\n<li><strong>8: Nodule/Mass + 13: Pulmonary fibrosis (82)</strong></li>\n<li>5: <strong>ILD + 6: Infiltration (75)</strong></li>\n<li>2: Calcification + 11: Pleural thickening (69)</li>\n<li>7: Lung Opacity + 11: Pleural thickening (60)</li>\n<li>5: ILD + 13: Pulmonary fibrosis (57)</li>\n<li>1: Atelectasis + 13: Pulmonary fibrosis (53)</li>\n<li>4: Consolidation + 8: Nodule/Mass (46)</li>\n<li><strong>5: ILD + 9: Other lesion (38)</strong></li>\n<li>2: Calcification + 9: Other lesion (38)</li>\n<li>2: Calcification + 8: Nodule/Mass (35)</li>\n<li>7: Lung Opacity + 10: Pleural effusion (30)</li>\n<li>11: Pleural thickening + 13: Pulmonary fibrosis (28)</li>\n<li>4: Consolidation + 7: Lung Opacity + 8: Nodule/Mass (25)</li>\n<li><strong>6: Infiltration + 7: Lung Opacity + 13: Pulmonary fibrosis (23)</strong></li>\n<li><strong>5: ILD + 9: Other lesion + 13: Pulmonary fibrosis (22)</strong></li>\n<li>1: Atelectasis + 7: Lung Opacity (20)</li>\n<li>9: Other lesion + 11: Pleural thickening (19)</li>\n<li>4: Consolidation + 13: Pulmonary fibrosis (19)</li>\n<li>1: Atelectasis + 4: Consolidation (18)</li>\n<li><em>0: Aortic enlargement + 2: Calcification (14)</em></li>\n<li><strong>5: ILD + 6: Infiltration + 13: Pulmonary fibrosis (14)</strong></li>\n<li><strong>4: Consolidation + 6: Infiltration + 13: Pulmonary fibrosis (14)</strong></li>\n<li><strong>6: Infiltration + 9: Other lesion + 13: Pulmonary fibrosis (10)</strong></li>\n<li>4: Consolidation + 6: Infiltration + 7: Lung Opacity (10)</li>\n</ul>\n<p>A few of the rarer combinations seem like they might well be wrong, e.g. \"0: Aortic enlargement + 3: Cardiomegaly\" for the same bounding box. Also, since cardiomegaly tends to be a really wide area and nodule/mass really small, something like \"3: Cardiomegaly + 8: Nodule/Mass\" does not seem so likely.</p>",
          "rawMarkdown": "I've [summarized those cases](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes#Bounding-boxes-with-more-than-one-label) now and I guess I'm not sure whether you would call all of those \"dense lesion area\"(s)?  I'm not saying that these don't mostly make sense, but rather that there might be other patterns (there's only very few, were I'm semi-confident to say they are wrong).\n\nE.g. I'm not sure whether one would describe the most common combination of pleural effusion and thickening as lesions (perhaps?)? Something like *0: Aortic enlargement + 2: Calcification* also does not exactly fit the pattern they stated (or would you call calcification a \"lesion\"?). I guess the cases below in **bold** that are more \"something you see\" (quite possibly a dense lesion area) + \"assigned disease diagnosis\" (such as pulmonary fibrosis or ILD) arguably fit the pattern (just the dense lesion area gets assigned a disease diagnosis, too).\n\nThe below are all the combinations that you have for at least 10 bounding boxes (there's obviously tons of rare combinations that happen for just 1), the number in brackets is the number of bounding boxes with this combination of labels:\n* 10: Pleural effusion + 11: Pleural thickening (1285)\n* **6: Infiltration + 13: Pulmonary fibrosis (318)**\n* 6: Infiltration + 7: Lung Opacity (216)\n* 7: Lung Opacity + 8: Nodule/Mass (161)\n* 7: Lung Opacity + 9: Other lesion (158)\n* 4: Consolidation + 7: Lung Opacity (155)\n* **2: Calcification + 13: Pulmonary fibrosis (153)**\n* **7: Lung Opacity + 13: Pulmonary fibrosis (138)**\n* 8: Nodule/Mass + 9: Other lesion (136)\n* **9: Other lesion + 13: Pulmonary fibrosis (106)**\n* **8: Nodule/Mass + 13: Pulmonary fibrosis (82)**\n* 5: **ILD + 6: Infiltration (75)**\n* 2: Calcification + 11: Pleural thickening (69)\n* 7: Lung Opacity + 11: Pleural thickening (60)\n* 5: ILD + 13: Pulmonary fibrosis (57)\n* 1: Atelectasis + 13: Pulmonary fibrosis (53)\n* 4: Consolidation + 8: Nodule/Mass (46)\n* **5: ILD + 9: Other lesion (38)**\n* 2: Calcification + 9: Other lesion (38)\n* 2: Calcification + 8: Nodule/Mass (35)\n* 7: Lung Opacity + 10: Pleural effusion (30)\n* 11: Pleural thickening + 13: Pulmonary fibrosis (28)\n* 4: Consolidation + 7: Lung Opacity + 8: Nodule/Mass (25)\n* **6: Infiltration + 7: Lung Opacity + 13: Pulmonary fibrosis (23)**\n* **5: ILD + 9: Other lesion + 13: Pulmonary fibrosis (22)**\n* 1: Atelectasis + 7: Lung Opacity (20)\n* 9: Other lesion + 11: Pleural thickening (19)\n* 4: Consolidation + 13: Pulmonary fibrosis (19)\n* 1: Atelectasis + 4: Consolidation (18)\n* *0: Aortic enlargement + 2: Calcification (14)*\n* **5: ILD + 6: Infiltration + 13: Pulmonary fibrosis (14)**\n* **4: Consolidation + 6: Infiltration + 13: Pulmonary fibrosis (14)**\n* **6: Infiltration + 9: Other lesion + 13: Pulmonary fibrosis (10)**\n* 4: Consolidation + 6: Infiltration + 7: Lung Opacity (10)\n\nA few of the rarer combinations seem like they might well be wrong, e.g. \"0: Aortic enlargement + 3: Cardiomegaly\" for the same bounding box. Also, since cardiomegaly tends to be a really wide area and nodule/mass really small, something like \"3: Cardiomegaly + 8: Nodule/Mass\" does not seem so likely.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1162066,
      "postDate": "2021-01-21T00:03:17.417Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1211810,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2021-02-20T15:26:02.103000",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>,<br>\nradiologist here. It is also challenging for a radiologist =) </p>\n<p>To clean cardiomegaly and aortic enlargement is pretty straigtforward and easy job, but there are some classes that would make pretty huge inter-observer variance:  ILD vs pulmonary fibrosis… or consolidation / infiltration / opacity. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1162125,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-21T02:23:01.440000",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> nice finding! I was just about to investigate this too, but without using any machine learning. When you look at bounding box centers for a class of findings - say Aortic Enlargement - and then check for bounding box centers that are 3 or more standard deviations from the mean, you find the potential errors pretty quickly. My <a href=\"https://www.kaggle.com/craigmthomas/localization-of-findings\" target=\"_blank\">quick EDA</a> does this. It shows the density heatmap next to the bounding box centers. On the bounding box center plot it measures how many standard deviations from the mean each bounding box center appears. This is for Aortic Enlargement:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2F70b393afd95cd06ce5d2f1efea60c87b%2Faortic_enlargment_bbox_centers.png?generation=1611195204811621&amp;alt=media\" alt=\"\"></p>\n<p>The red dots trailing down near the center are more than 4 standard deviations from the mean. I took a quick look at the actual bounding box for some of them, and my initial thoughts were that they must have been Cardiomegaly that were accidentally mislabeled as Aortic Enlargement (but I'm not a radiologist, so I can't say that definitively). Your machine learning approach for detection is pretty neat. Thanks for sharing!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1211813,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2021-02-20T15:29:35.243000",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a>,</p>\n<p>nice visualisation there!<br>\nThe problem with the absolute position and it's deviation from the mean is that sometimes (not to rarely) the positioning of the patient i off-center, causing a shift of all organ projections on the image.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1211979,
          "author_name": "Craig Thomas",
          "author_url": "",
          "post_date": "2021-02-20T18:05:32.120000",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> thank you, your comment is very helpful! Being a radiologist, I'm sure you've definitely seen a lot of different x-ray alignments. In fact, perhaps you may be able to give us an idea about how many rotational degrees we could apply to our training set to achieve a realistic set of transforms to bolster our data (i.e. a rotation of 2 - 5 degrees may be appropriate, but 180 degrees is not likely to occur)? As well, maybe insights into the X and Y skews that may be appropriate?</p>\n<p>As to my visualization, I should have prefaced that finding with the fact that most of those &lt; 1% occurrences should be investigated manually to see if they are misclassifications (definitely have to be careful with applying any kind of automated cleaning). Most certainly, deviation from the mean may not be the best way to detect outliers in this context, but box centers occurring more than 4 standard deviations are worth investigating. With that said, I did that for Aortic Enlargment and found that some of the lower red outliers look very much like boxes for Cardiomegaly. There were also some higher up that looked like Pleural Thickening. </p>\n<p>As per your comment, I found the utility of doing this outside of Aortic Enlargement and Cardiomegaly is limited. Both Aortic Enlargement and Cardiomegaly tend to be very tightly grouped in the test data. Certainly when you get into Pleural Effusion and other classes, absolute location isn't a reliable indicator of mislabelled boxes, and random shifts in patient alignment have a much greater impact on what could potentially be an outlier. </p>\n<p>Thanks again for the comment, I really appreciate it!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212695,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2021-02-21T13:44:37.597000",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a> ,<br>\nthe question of the augmentation is very interesting and not easily to answer. By the chest x-rays the \"rotation around the midpoint of the image\" occurs rather in only a few degrees range, since this is something what the assistant sees and controls during the acquisition! More problematic is the rotation of the patient of his own axis and the projection's center point ( imagine moving the radiation source in y axis as seen on the second image). Both of them changes the projection of all organs and the appearance of the pathologies and normal anatomy to.</p>\n<p><img src=\"https://i.postimg.cc/rpf5NW4m/xray4-1.jpg\" alt=\"chest\"> </p>\n<p>There are cases where the images are upside down or rotated to x*π/2 =)</p>\n<p>Nono, the method is VERY good, in fact i would do something similar to select outliers to look for for the first time cleaning. I know that there are many false labelings in the set. Maybe i find the time and annotate a random 5k aortic knob (from the non enlarged ones), so one will be able to train a \"aortic knob detection\", would be a good accompanying set to <a href=\"https://www.kaggle.com/sandorkonya/5k-trachea-bifurcation-on-chest-xray\" target=\"_blank\">this one</a>.</p>\n<p>Cooccurence of cardiovascular diseases is a known fact, nice to see the theory in practice here =)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1213313,
          "author_name": "Craig Thomas",
          "author_url": "",
          "post_date": "2021-02-22T03:42:25.260000",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> thank you so much for your comment! Your insight about the projection angle changing both normal and pathological appearance is very interesting. As well, the idea of an aortic knob detector is also quite interesting. Having the ability to create a waypoint like that on the image may help. I've been working on an ensemble solution, one where there is a classifier and bounding box regressor for each class of finding. Each classifier and regressor tries to focus on specific areas of the x-ray. The classifier portion works well for aortic enlargement and cardiomegaly because they tend to appear in roughly the same absolute area (the regressors are more of a challenge due to ground truth). Overall just focusing on these two classes has yielded a LB score of 0.08, however, identifying features to use purely for positioning purposes could potentially help out a lot - they could be used zoom in on the target areas more accurately. Thanks again for your comments, I really appreciate them!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1217842,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2021-02-25T11:09:33.047000",
          "content": "<p>The idea with the regressor / classifier is good, but as I was browsing the labelled images, I found several classes to have pretty arbitrarily placed ROIs…<br>\nEspecially in case of ILD &amp;Fibrosis - these are the 2 where one can argue why they chose especially the position for that (see for example your 6th image).<br>\nIn case of these 2 labels it would be the best not to take the labels of the GT but the whole lung area (like in the second example), and create a classifier for the disease and then analyze the distribution of GT for the most likely position &amp; size for these bounding boxes. As you see on the second example, the classes ILD &amp; Fibrosis are not even separated… cooccur on the same bbox.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1220776,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2021-02-28T10:55:31.787000",
          "content": "<p><a href=\"https://www.kaggle.com/craigmthomas\" target=\"_blank\">@craigmthomas</a>,</p>\n<p>FYI <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222626\" target=\"_blank\">this</a>.<br>\nExactly this is the way how the pathologies also \"change\" their appearance depending on the projection. Nevertheless this is a problem for mainly the longitudinal data analysis (in the current competition we have no duplicate of patients) but it is surely a factor that influences the inter-observer  error.</p>\n<p>I reviewed (…okay, that's to much…  i just scrolled through) the 15k training data and there are ~few dozens of images with an appearance of a portable/bedside image so in this competition this kind of bias play rather a smaller role!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1161399,
      "author_name": "T Nakamura",
      "author_url": "",
      "post_date": "2021-01-20T14:49:36.913000",
      "content": "<p>I found another annotation problem.<br>\nMultiple labels are assigned to exact same bbox, this can be found in 1775 images (mostly annotated by R8, R9, R10).<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1111209%2F43e93cd1401dfd1df532fdd1f37473de%2Fss.jpg?generation=1611154145439806&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1161434,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-20T15:10:48.727000",
          "content": "<p>Interesting. I wonder whether that is actually something that makes sense or is a data issues… E.g. infiltration and pulmonary fibrosis co-occur a lot (correlation 0.38), so if the infiltration is what the radiologist thinks is what identifies the pulmonary fibrosis, then maybe this is sensible?!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1161537,
          "author_name": "Geir Drange",
          "author_url": "",
          "post_date": "2021-01-20T15:57:45.503000",
          "content": "<p>Is it a problem? This is also known as multi-label classification.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161564,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-20T16:10:31.460000",
          "content": "<p>Not a problem, if it's correct…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161967,
          "author_name": "LAZCoder",
          "author_url": "",
          "post_date": "2021-01-20T21:24:15.193000",
          "content": "<p>This should not be a problem with annotations as stated also from competition host in <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/209298\" target=\"_blank\">this thread</a></p>\n<pre><code>Some dense lesion area contains many labels, our radiologists created a box then assign many labels to that box.\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1162067,
          "author_name": "T Nakamura",
          "author_url": "",
          "post_date": "2021-01-21T00:04:39.437000",
          "content": "<p><a href=\"https://www.kaggle.com/lazcoder\" target=\"_blank\">@lazcoder</a> Thank you ! I missed the thread.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162617,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-21T08:36:58.240000",
          "content": "<p>I've <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes#Bounding-boxes-with-more-than-one-label\" target=\"_blank\">summarized those cases</a> now and I guess I'm not sure whether you would call all of those \"dense lesion area\"(s)?  I'm not saying that these don't mostly make sense, but rather that there might be other patterns (there's only very few, were I'm semi-confident to say they are wrong).</p>\n<p>E.g. I'm not sure whether one would describe the most common combination of pleural effusion and thickening as lesions (perhaps?)? Something like <em>0: Aortic enlargement + 2: Calcification</em> also does not exactly fit the pattern they stated (or would you call calcification a \"lesion\"?). I guess the cases below in <strong>bold</strong> that are more \"something you see\" (quite possibly a dense lesion area) + \"assigned disease diagnosis\" (such as pulmonary fibrosis or ILD) arguably fit the pattern (just the dense lesion area gets assigned a disease diagnosis, too).</p>\n<p>The below are all the combinations that you have for at least 10 bounding boxes (there's obviously tons of rare combinations that happen for just 1), the number in brackets is the number of bounding boxes with this combination of labels:</p>\n<ul>\n<li>10: Pleural effusion + 11: Pleural thickening (1285)</li>\n<li><strong>6: Infiltration + 13: Pulmonary fibrosis (318)</strong></li>\n<li>6: Infiltration + 7: Lung Opacity (216)</li>\n<li>7: Lung Opacity + 8: Nodule/Mass (161)</li>\n<li>7: Lung Opacity + 9: Other lesion (158)</li>\n<li>4: Consolidation + 7: Lung Opacity (155)</li>\n<li><strong>2: Calcification + 13: Pulmonary fibrosis (153)</strong></li>\n<li><strong>7: Lung Opacity + 13: Pulmonary fibrosis (138)</strong></li>\n<li>8: Nodule/Mass + 9: Other lesion (136)</li>\n<li><strong>9: Other lesion + 13: Pulmonary fibrosis (106)</strong></li>\n<li><strong>8: Nodule/Mass + 13: Pulmonary fibrosis (82)</strong></li>\n<li>5: <strong>ILD + 6: Infiltration (75)</strong></li>\n<li>2: Calcification + 11: Pleural thickening (69)</li>\n<li>7: Lung Opacity + 11: Pleural thickening (60)</li>\n<li>5: ILD + 13: Pulmonary fibrosis (57)</li>\n<li>1: Atelectasis + 13: Pulmonary fibrosis (53)</li>\n<li>4: Consolidation + 8: Nodule/Mass (46)</li>\n<li><strong>5: ILD + 9: Other lesion (38)</strong></li>\n<li>2: Calcification + 9: Other lesion (38)</li>\n<li>2: Calcification + 8: Nodule/Mass (35)</li>\n<li>7: Lung Opacity + 10: Pleural effusion (30)</li>\n<li>11: Pleural thickening + 13: Pulmonary fibrosis (28)</li>\n<li>4: Consolidation + 7: Lung Opacity + 8: Nodule/Mass (25)</li>\n<li><strong>6: Infiltration + 7: Lung Opacity + 13: Pulmonary fibrosis (23)</strong></li>\n<li><strong>5: ILD + 9: Other lesion + 13: Pulmonary fibrosis (22)</strong></li>\n<li>1: Atelectasis + 7: Lung Opacity (20)</li>\n<li>9: Other lesion + 11: Pleural thickening (19)</li>\n<li>4: Consolidation + 13: Pulmonary fibrosis (19)</li>\n<li>1: Atelectasis + 4: Consolidation (18)</li>\n<li><em>0: Aortic enlargement + 2: Calcification (14)</em></li>\n<li><strong>5: ILD + 6: Infiltration + 13: Pulmonary fibrosis (14)</strong></li>\n<li><strong>4: Consolidation + 6: Infiltration + 13: Pulmonary fibrosis (14)</strong></li>\n<li><strong>6: Infiltration + 9: Other lesion + 13: Pulmonary fibrosis (10)</strong></li>\n<li>4: Consolidation + 6: Infiltration + 7: Lung Opacity (10)</li>\n</ul>\n<p>A few of the rarer combinations seem like they might well be wrong, e.g. \"0: Aortic enlargement + 3: Cardiomegaly\" for the same bounding box. Also, since cardiomegaly tends to be a really wide area and nodule/mass really small, something like \"3: Cardiomegaly + 8: Nodule/Mass\" does not seem so likely.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1162066,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-21T00:03:17.417000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1161322": "Based on some of the [discussions](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/211035), where we uncovered some mislabelled bounding boxes, that suggested that at least some 'aortic enlargement' and 'cardiomegaly' labels for bounding boxes are wrong, I realized that there's definitely some mislabeled bounding boxes.\n\nShould we care? I think so, if the training data is not generated from the same process as the test data (2 additional radiologists re-conciliate issues for the test data) and a lot of the differences will be mistakes in the training data that would likely be fixed, if the image were in the test data. Thus, it makes sense to try and fix the issues we can fix on the training data (that should improve the link from training and CV performance to LB).\n\nThus, I tried one approach similar to how we uncovered issues in the forum discussions by using LightGBM on the `train.csv` and some additional `.dicom` meta-data. **The notebook implementing this model and producing the images below is [here](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes/).**\n\nHere are some examples of issues that I found. I'm not a radiologist, so take my human assessment of these with a grain of salt (expert input very, very welcome!). Firstly, the approach again picked up some of those  'aortic enlargement' vs. 'cardiomegaly' issues:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F2493608399172fec8083d7a821336ffc%2F__results___41_3.png?generation=1611149547487964&alt=media)\nI'm also pretty sure that cardiomegaly should not be in this location, so presumably the model is right that this could be aortic enlargement?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fc237c3c34fdac9358d347ce07d4883cb%2F__results___41_9.png?generation=1611149947019757&alt=media)\nHere's a case of aortic enlargement labelled in a weird position (the one on the left):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fa9738147a42ecb7d6e7a54c5a5b82ddf%2F__results___41_5.png?generation=1611149609228968&alt=media)\nHere's another example of a weird position for a class and very small bounding box by the standards of the class (so the aortic enlargement suggestion of the model does make sense to me):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F44dc74f173eacdeff9501f741a713b7a%2F__results___41_10.png?generation=1611150122673016&alt=media)\nBut sometimes, I just don't know what to make of it. E.g. is the model right that this is pleural thickening rather than pneumothorax ([EDA suggests](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#5.-How-big-do-bounding-boxes-tend-to-be-for-different-classes?-How-many-are-there?) that pneumothorax tends to be very large bounding boxes and pleural thickening tends to be in [this particular region](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?) of x-rays with smaller bounding boxes, but I cannot really judge this well)? \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fee3d330f108eb32b749b591b964f33de%2F__results___41_8.png?generation=1611149721332674&alt=media)\nI've identified at least one case, where I think my model is wrong (in the sense that it looks like the radiologist had already labelled an aortic enlargement and clearly did intend to label the overlapping - but not identical - bounding box as Nodule/Mass) - perhaps I could avoid this (presumably) wrong prediction by adding a feature counting other labels assigned by the same radiologist to other bounding boxes (**Update:** Adding the extra feature that I mentioned in version 17 of the notebook seems to have stopped the model from capturing the image I thought it should not identify):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2Fec64766573aa794fee39c841d7a0caa9%2F__results___41_46.png?generation=1611150416122729&alt=media)\nThe second (higher) box here seems wrong, but who knows whether the model is right that this is ILD:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F1f493c5d1c3d677038122836e748e839%2F__results___41_52.png?generation=1611150671637265&alt=media)\n\n\nThere's plenty more examples in the notebook, but a lot of the cases seem to involve aortic enlargement, cardiomegaly, pleural effusion/thickening and occasionally a few other classes.\n\nIn any case, it looks like it might be worthwhile to try to put together a cleaned up dataset, but it's challenging as a non-radiologist.",
    "1211810": "@bjoernholzhauer,\nradiologist here. It is also challenging for a radiologist =) \n\nTo clean cardiomegaly and aortic enlargement is pretty straigtforward and easy job, but there are some classes that would make pretty huge inter-observer variance:  ILD vs pulmonary fibrosis... or consolidation / infiltration / opacity. \n",
    "1162125": "@bjoernholzhauer nice finding! I was just about to investigate this too, but without using any machine learning. When you look at bounding box centers for a class of findings - say Aortic Enlargement - and then check for bounding box centers that are 3 or more standard deviations from the mean, you find the potential errors pretty quickly. My [quick EDA](https://www.kaggle.com/craigmthomas/localization-of-findings) does this. It shows the density heatmap next to the bounding box centers. On the bounding box center plot it measures how many standard deviations from the mean each bounding box center appears. This is for Aortic Enlargement:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2051313%2F70b393afd95cd06ce5d2f1efea60c87b%2Faortic_enlargment_bbox_centers.png?generation=1611195204811621&alt=media)\n\nThe red dots trailing down near the center are more than 4 standard deviations from the mean. I took a quick look at the actual bounding box for some of them, and my initial thoughts were that they must have been Cardiomegaly that were accidentally mislabeled as Aortic Enlargement (but I'm not a radiologist, so I can't say that definitively). Your machine learning approach for detection is pretty neat. Thanks for sharing!",
    "1161399": "I found another annotation problem.\nMultiple labels are assigned to exact same bbox, this can be found in 1775 images (mostly annotated by R8, R9, R10).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1111209%2F43e93cd1401dfd1df532fdd1f37473de%2Fss.jpg?generation=1611154145439806&alt=media)",
    "1162066": ""
  }
}