{
  "id": 229642,
  "title": "Data Labeling Issues - Structured Bias",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229642",
  "author_name": "beluga",
  "post_date": "2021-03-31T04:34:13.709000",
  "votes": 27,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First, I would like to thank for the organizers for collecting the data and hosting this interesting challenge. It was my first kaggle competition with medical images and I learned a lot, special thanks for my teammates ( <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> for object detection and <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for radiology) for that!<br>\nCongrats to the winners, I am really curious how they cracked this one and I am waiting for their summaries.</p>\n<p>Meanwhile I would like share a few serious bias/data leakage/structured noise issues we found that could hurt any real world application built on this dataset.  </p>\n<h2>Previous External Data Issues</h2>\n<p>I started the competition by reviewing the organizers paper and the previous external chest X-ray datasets (CRX14, Chexpert) That way I found a few interesting blog posts about the possible problems with the already published datasets.</p>\n<ul>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2017/11/18/quick-thoughts-on-chestxray14-performance-claims-and-clinical-tasks/\" target=\"_blank\">Quick thoughts on ChestXray14, performance claims, and clinical tasks</a></li>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\" target=\"_blank\">Exploring the ChestXray14 dataset: problems</a></li>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2019/02/25/half-a-million-x-rays-first-impressions-of-the-stanford-and-mit-chest-x-ray-datasets/\" target=\"_blank\">Half a million x-rays! First impressions of the Stanford and MIT chest x-ray datasets</a></li>\n</ul>\n<p>The organizers improved some of the already mentioned problems (NLP labels, image quality, AP/PA Xrays, multiple images per patient) but unfortunately they also introduced others.</p>\n<h2>DICOM meta data</h2>\n<p>First we checked the DICOM meta data and extracted all the fields we could.<br>\nWe trained a few XGB binary models for each class and surprisingly they gave <strong>pretty good local validation score without using deep learning or even looking at the actual images!</strong></p>\n<p><img src=\"https://storage.googleapis.com/vinbigdata/XGB.png\" alt=\"\"></p>\n<p>That's weird! It means we have a serious leakage/bias in these fields.<br>\nSome of the most important features were Age, Sex and <strong>SourceApplicationEntityTitle (SAET)</strong>. <br>\nWhile Age and Sex could be important in medical diagnostics I can not explain the following normality rates by Sex:</p>\n<ul>\n<li>F: 15%</li>\n<li>M: 17%</li>\n<li>O: 61%</li>\n</ul>\n<h2>No Finding class</h2>\n<p>It was discussed in the forums that the No Finding class had 100% agreement rate. That is very unusual in human annotation processes especially for such difficult and ambigous task.<br>\nIt also means that the majority of the training set is not really annotated.</p>\n<h2>Annotator Group – SourceApplicationTitle</h2>\n<p>Analyzing the radiologist groups also lead to strange patterns R8-R9-R10  and R11-R17 always annotated together. R1-R7 always (except maybe one case) annotated No Finding.<br>\nIf we check the annotations by SAET the pattern is a bit more clear<br>\n<img src=\"https://storage.googleapis.com/vinbigdata/saet_rad.png\" alt=\"\"></p>\n<p><strong>The VITREA1 images did not have a single positive label as they were \"annotated\" by R1-R7.</strong><br>\nUnfortunately these images took 53% of the test set. The vast majority of the positive training images came with Unknown SAET and were annotated by R8-R10, and unfortunately these images were 26% of the test set.</p>\n<p>I found the fact that many kagglers reported binary Normal/Abnormal models with AUC 0.992 very weird. We made a few tests and were able to learn the SAET almost perfectly. This kind of structured label bias could be very dangerous in real world medical application.</p>\n<h2>Class ambiguity</h2>\n<p>This one is the least serious one. I think at this point it is difficult to change too much on the pathology classes without loosing the ability to utilize the already collected hundreds of thousands public x-rays. It has been already discussed that the visual difference between classes could be quite difficult to define especially for labels consolidation / infiltration / atelectasis / pneumonia. </p>\n<p>Just another example, I wanted to understand better the ILD class, my quick google search resulted:</p>\n<p><em>\"<strong>Interstitial lung disease (ILD)</strong> is another term for <strong>pulmonary fibrosis</strong>, which means “scarring” and “inflammation” of the interstitium (the tissue that surrounds the lung's air sacs, blood vessels and airways). This scarring makes the lung tissue stiff, which can make breathing difficult.Apr 26, 2018\"</em></p>\n<p>In a perfect world where annotators agree in 100% time each pathology would be found by all three annotators. Of course that is impossible even for easier annotation tasks (Chihuahua or muffin). On the training data the most obvious classes (Cardiomegaly, Aortic Enlargement) had perfect agreement 56-57% of the time while the most difficult classes (Consolidation, Other Lesion, Atelectasis) had only 10% triple agreement. That is quite a lot label noise to deal with and we have only 4K positive images in the training set.</p>\n<h2>Bounding box habits</h2>\n<p>Object detection is much more complex than classification. It is also true for collecting ground truth labels.  Even if we decreased the minimum IOU to 0.4 lots of the annotated boxes did not agree at all. There were systematic differences between the radiologist groups too. R8-R10 usually used much smaller boxes and they were more consistent, while R11+ used less precise boxes. <br>\n<img src=\"https://storage.googleapis.com/vinbigdata/AortaAnnotation.png\" alt=\"\"></p>\n<p>Another systematic difference was that quite often R8-R10 only annotated the lower part of the heart while R11-R17 used larger boxes even bigger than the full heart.</p>\n<p><img src=\"https://storage.googleapis.com/vinbigdata/Cardio.png\" alt=\"\"><br>\nAs we have seen it is possible to learn these patterns with enough data but it does not make sense for real world applications.</p>\n<h2>Train-Test labeling differences</h2>\n<p>I think using different train and test labeling method made the task even more difficult. We have already seen, that the the raw radiologist annotations had quite significant label noise. Using two unknown senior radiologists to clean the test set is not enough. Again, they could have different labeling habits (e.g. many precise small boxes, just one big box the lung section, lazy default box fusion etc.) and even keep/introduce more structured bias. </p>\n<p>I don't know how much would it cost to annotate the full training set as accurately as the test was labeled (My estimation is 1 min per image ~ 300 hours per radiologist) If that is too much, than providing 3K validation set annotated by the same procedure as the test set should be feasible. </p>\n<p>These kind of train-test labeling differences lead to LB probing/LB overfitting and do not actually contribute to better models.</p>\n<p>Thanks to all these issues, it is quite easy to train a simple baseline yolov5 model wtih 0.9+ AP local cross-validation for Aortic enlargement/Cardiomegaly then try it on the LB and realize it only achieves 0.15-0.4 AP. </p>\n<p>I have a few ideas that could explain some of the above issues (e.g. xrays coming from different hospitals/machines etc.) but I wonder how the host or other competitiors see it…</p>",
  "messages": [
    {
      "id": 1257745,
      "postDate": "2021-03-31T04:34:13.710Z",
      "content": "<p>First, I would like to thank for the organizers for collecting the data and hosting this interesting challenge. It was my first kaggle competition with medical images and I learned a lot, special thanks for my teammates ( <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> for object detection and <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for radiology) for that!<br>\nCongrats to the winners, I am really curious how they cracked this one and I am waiting for their summaries.</p>\n<p>Meanwhile I would like share a few serious bias/data leakage/structured noise issues we found that could hurt any real world application built on this dataset.  </p>\n<h2>Previous External Data Issues</h2>\n<p>I started the competition by reviewing the organizers paper and the previous external chest X-ray datasets (CRX14, Chexpert) That way I found a few interesting blog posts about the possible problems with the already published datasets.</p>\n<ul>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2017/11/18/quick-thoughts-on-chestxray14-performance-claims-and-clinical-tasks/\" target=\"_blank\">Quick thoughts on ChestXray14, performance claims, and clinical tasks</a></li>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\" target=\"_blank\">Exploring the ChestXray14 dataset: problems</a></li>\n<li><a href=\"https://lukeoakdenrayner.wordpress.com/2019/02/25/half-a-million-x-rays-first-impressions-of-the-stanford-and-mit-chest-x-ray-datasets/\" target=\"_blank\">Half a million x-rays! First impressions of the Stanford and MIT chest x-ray datasets</a></li>\n</ul>\n<p>The organizers improved some of the already mentioned problems (NLP labels, image quality, AP/PA Xrays, multiple images per patient) but unfortunately they also introduced others.</p>\n<h2>DICOM meta data</h2>\n<p>First we checked the DICOM meta data and extracted all the fields we could.<br>\nWe trained a few XGB binary models for each class and surprisingly they gave <strong>pretty good local validation score without using deep learning or even looking at the actual images!</strong></p>\n<p><img src=\"https://storage.googleapis.com/vinbigdata/XGB.png\" alt=\"\"></p>\n<p>That's weird! It means we have a serious leakage/bias in these fields.<br>\nSome of the most important features were Age, Sex and <strong>SourceApplicationEntityTitle (SAET)</strong>. <br>\nWhile Age and Sex could be important in medical diagnostics I can not explain the following normality rates by Sex:</p>\n<ul>\n<li>F: 15%</li>\n<li>M: 17%</li>\n<li>O: 61%</li>\n</ul>\n<h2>No Finding class</h2>\n<p>It was discussed in the forums that the No Finding class had 100% agreement rate. That is very unusual in human annotation processes especially for such difficult and ambigous task.<br>\nIt also means that the majority of the training set is not really annotated.</p>\n<h2>Annotator Group – SourceApplicationTitle</h2>\n<p>Analyzing the radiologist groups also lead to strange patterns R8-R9-R10  and R11-R17 always annotated together. R1-R7 always (except maybe one case) annotated No Finding.<br>\nIf we check the annotations by SAET the pattern is a bit more clear<br>\n<img src=\"https://storage.googleapis.com/vinbigdata/saet_rad.png\" alt=\"\"></p>\n<p><strong>The VITREA1 images did not have a single positive label as they were \"annotated\" by R1-R7.</strong><br>\nUnfortunately these images took 53% of the test set. The vast majority of the positive training images came with Unknown SAET and were annotated by R8-R10, and unfortunately these images were 26% of the test set.</p>\n<p>I found the fact that many kagglers reported binary Normal/Abnormal models with AUC 0.992 very weird. We made a few tests and were able to learn the SAET almost perfectly. This kind of structured label bias could be very dangerous in real world medical application.</p>\n<h2>Class ambiguity</h2>\n<p>This one is the least serious one. I think at this point it is difficult to change too much on the pathology classes without loosing the ability to utilize the already collected hundreds of thousands public x-rays. It has been already discussed that the visual difference between classes could be quite difficult to define especially for labels consolidation / infiltration / atelectasis / pneumonia. </p>\n<p>Just another example, I wanted to understand better the ILD class, my quick google search resulted:</p>\n<p><em>\"<strong>Interstitial lung disease (ILD)</strong> is another term for <strong>pulmonary fibrosis</strong>, which means “scarring” and “inflammation” of the interstitium (the tissue that surrounds the lung's air sacs, blood vessels and airways). This scarring makes the lung tissue stiff, which can make breathing difficult.Apr 26, 2018\"</em></p>\n<p>In a perfect world where annotators agree in 100% time each pathology would be found by all three annotators. Of course that is impossible even for easier annotation tasks (Chihuahua or muffin). On the training data the most obvious classes (Cardiomegaly, Aortic Enlargement) had perfect agreement 56-57% of the time while the most difficult classes (Consolidation, Other Lesion, Atelectasis) had only 10% triple agreement. That is quite a lot label noise to deal with and we have only 4K positive images in the training set.</p>\n<h2>Bounding box habits</h2>\n<p>Object detection is much more complex than classification. It is also true for collecting ground truth labels.  Even if we decreased the minimum IOU to 0.4 lots of the annotated boxes did not agree at all. There were systematic differences between the radiologist groups too. R8-R10 usually used much smaller boxes and they were more consistent, while R11+ used less precise boxes. <br>\n<img src=\"https://storage.googleapis.com/vinbigdata/AortaAnnotation.png\" alt=\"\"></p>\n<p>Another systematic difference was that quite often R8-R10 only annotated the lower part of the heart while R11-R17 used larger boxes even bigger than the full heart.</p>\n<p><img src=\"https://storage.googleapis.com/vinbigdata/Cardio.png\" alt=\"\"><br>\nAs we have seen it is possible to learn these patterns with enough data but it does not make sense for real world applications.</p>\n<h2>Train-Test labeling differences</h2>\n<p>I think using different train and test labeling method made the task even more difficult. We have already seen, that the the raw radiologist annotations had quite significant label noise. Using two unknown senior radiologists to clean the test set is not enough. Again, they could have different labeling habits (e.g. many precise small boxes, just one big box the lung section, lazy default box fusion etc.) and even keep/introduce more structured bias. </p>\n<p>I don't know how much would it cost to annotate the full training set as accurately as the test was labeled (My estimation is 1 min per image ~ 300 hours per radiologist) If that is too much, than providing 3K validation set annotated by the same procedure as the test set should be feasible. </p>\n<p>These kind of train-test labeling differences lead to LB probing/LB overfitting and do not actually contribute to better models.</p>\n<p>Thanks to all these issues, it is quite easy to train a simple baseline yolov5 model wtih 0.9+ AP local cross-validation for Aortic enlargement/Cardiomegaly then try it on the LB and realize it only achieves 0.15-0.4 AP. </p>\n<p>I have a few ideas that could explain some of the above issues (e.g. xrays coming from different hospitals/machines etc.) but I wonder how the host or other competitiors see it…</p>",
      "rawMarkdown": "First, I would like to thank for the organizers for collecting the data and hosting this interesting challenge. It was my first kaggle competition with medical images and I learned a lot, special thanks for my teammates ( @pestipeti for object detection and @sandorkonya for radiology) for that!\nCongrats to the winners, I am really curious how they cracked this one and I am waiting for their summaries.\n\n\nMeanwhile I would like share a few serious bias/data leakage/structured noise issues we found that could hurt any real world application built on this dataset.  \n\n## Previous External Data Issues\nI started the competition by reviewing the organizers paper and the previous external chest X-ray datasets (CRX14, Chexpert) That way I found a few interesting blog posts about the possible problems with the already published datasets.\n\n* [Quick thoughts on ChestXray14, performance claims, and clinical tasks](https://lukeoakdenrayner.wordpress.com/2017/11/18/quick-thoughts-on-chestxray14-performance-claims-and-clinical-tasks/)\n* [Exploring the ChestXray14 dataset: problems](https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/)\n* [Half a million x-rays! First impressions of the Stanford and MIT chest x-ray datasets](https://lukeoakdenrayner.wordpress.com/2019/02/25/half-a-million-x-rays-first-impressions-of-the-stanford-and-mit-chest-x-ray-datasets/)\n\nThe organizers improved some of the already mentioned problems (NLP labels, image quality, AP/PA Xrays, multiple images per patient) but unfortunately they also introduced others.\n\n\n## DICOM meta data\n\nFirst we checked the DICOM meta data and extracted all the fields we could.\nWe trained a few XGB binary models for each class and surprisingly they gave **pretty good local validation score without using deep learning or even looking at the actual images!**\n\n![](https://storage.googleapis.com/vinbigdata/XGB.png)\n\nThat's weird! It means we have a serious leakage/bias in these fields.\nSome of the most important features were Age, Sex and **SourceApplicationEntityTitle (SAET)**. \nWhile Age and Sex could be important in medical diagnostics I can not explain the following normality rates by Sex:\n* F: 15%\n* M: 17%\n* O: 61%\n\n## No Finding class\nIt was discussed in the forums that the No Finding class had 100% agreement rate. That is very unusual in human annotation processes especially for such difficult and ambigous task.\nIt also means that the majority of the training set is not really annotated.\n\n## Annotator Group – SourceApplicationTitle\nAnalyzing the radiologist groups also lead to strange patterns R8-R9-R10  and R11-R17 always annotated together. R1-R7 always (except maybe one case) annotated No Finding.\nIf we check the annotations by SAET the pattern is a bit more clear\n![](https://storage.googleapis.com/vinbigdata/saet_rad.png)\n\n**The VITREA1 images did not have a single positive label as they were \"annotated\" by R1-R7.**\nUnfortunately these images took 53% of the test set. The vast majority of the positive training images came with Unknown SAET and were annotated by R8-R10, and unfortunately these images were 26% of the test set.\n\nI found the fact that many kagglers reported binary Normal/Abnormal models with AUC 0.992 very weird. We made a few tests and were able to learn the SAET almost perfectly. This kind of structured label bias could be very dangerous in real world medical application.\n\n## Class ambiguity\nThis one is the least serious one. I think at this point it is difficult to change too much on the pathology classes without loosing the ability to utilize the already collected hundreds of thousands public x-rays. It has been already discussed that the visual difference between classes could be quite difficult to define especially for labels consolidation / infiltration / atelectasis / pneumonia. \n\nJust another example, I wanted to understand better the ILD class, my quick google search resulted:\n\n*\"**Interstitial lung disease (ILD)** is another term for **pulmonary fibrosis**, which means “scarring” and “inflammation” of the interstitium (the tissue that surrounds the lung's air sacs, blood vessels and airways). This scarring makes the lung tissue stiff, which can make breathing difficult.Apr 26, 2018\"*\n\nIn a perfect world where annotators agree in 100% time each pathology would be found by all three annotators. Of course that is impossible even for easier annotation tasks (Chihuahua or muffin). On the training data the most obvious classes (Cardiomegaly, Aortic Enlargement) had perfect agreement 56-57% of the time while the most difficult classes (Consolidation, Other Lesion, Atelectasis) had only 10% triple agreement. That is quite a lot label noise to deal with and we have only 4K positive images in the training set.\n \n## Bounding box habits\n \nObject detection is much more complex than classification. It is also true for collecting ground truth labels.  Even if we decreased the minimum IOU to 0.4 lots of the annotated boxes did not agree at all. There were systematic differences between the radiologist groups too. R8-R10 usually used much smaller boxes and they were more consistent, while R11+ used less precise boxes. \n![](https://storage.googleapis.com/vinbigdata/AortaAnnotation.png)\n\nAnother systematic difference was that quite often R8-R10 only annotated the lower part of the heart while R11-R17 used larger boxes even bigger than the full heart.\n\n![](https://storage.googleapis.com/vinbigdata/Cardio.png)\nAs we have seen it is possible to learn these patterns with enough data but it does not make sense for real world applications.\n\n\n## Train-Test labeling differences\nI think using different train and test labeling method made the task even more difficult. We have already seen, that the the raw radiologist annotations had quite significant label noise. Using two unknown senior radiologists to clean the test set is not enough. Again, they could have different labeling habits (e.g. many precise small boxes, just one big box the lung section, lazy default box fusion etc.) and even keep/introduce more structured bias. \n\nI don't know how much would it cost to annotate the full training set as accurately as the test was labeled (My estimation is 1 min per image ~ 300 hours per radiologist) If that is too much, than providing 3K validation set annotated by the same procedure as the test set should be feasible. \n\nThese kind of train-test labeling differences lead to LB probing/LB overfitting and do not actually contribute to better models.\n\nThanks to all these issues, it is quite easy to train a simple baseline yolov5 model wtih 0.9+ AP local cross-validation for Aortic enlargement/Cardiomegaly then try it on the LB and realize it only achieves 0.15-0.4 AP. \n\nI have a few ideas that could explain some of the above issues (e.g. xrays coming from different hospitals/machines etc.) but I wonder how the host or other competitiors see it...\n",
      "votes": 27
    },
    {
      "id": 1257891,
      "postDate": "2021-03-31T07:26:56.213Z",
      "content": "<p>I fully agree with you <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> - there were some really weird things happening in this data.</p>\n<ul>\n<li><p>There was some extreme bias in what images which annotator annotates. I have the strong hypothesis that image distribution was not fully random. You could fit a classifier predicting the annotator for an image with very good accuracy. Actually we tried to remove the bias to better adjust to possible pattern changes in test, but it was basically impossible to remove the strong bias.</p></li>\n<li><p>The fact that each image either has three agreeing annotators for non-findings, or for findings, does not reflect any realistic use case. There should be images where some annotators say it is non-finding and others say it is a finding. There has to be some specific process in place.</p></li>\n<li><p>You have a strong difference between train and test labels, but you have no realistic chance of finding a really strong CV setup here. Public test is too small (300 images with large portion of nonfindings) to make too strong conclusions, and local CV is just some random effort to guess how test could look like. I think this setup is far from reality where you would build some realistic validation set. I can see the issue of having more \"badly\" labeled vs. \"well\" labeled data, but in this case having a specific validation dataset that is closer to test would produce way better solutions. </p></li>\n</ul>",
      "rawMarkdown": "I fully agree with you @gaborfodor - there were some really weird things happening in this data.\n\n- There was some extreme bias in what images which annotator annotates. I have the strong hypothesis that image distribution was not fully random. You could fit a classifier predicting the annotator for an image with very good accuracy. Actually we tried to remove the bias to better adjust to possible pattern changes in test, but it was basically impossible to remove the strong bias.\n\n- The fact that each image either has three agreeing annotators for non-findings, or for findings, does not reflect any realistic use case. There should be images where some annotators say it is non-finding and others say it is a finding. There has to be some specific process in place.\n\n- You have a strong difference between train and test labels, but you have no realistic chance of finding a really strong CV setup here. Public test is too small (300 images with large portion of nonfindings) to make too strong conclusions, and local CV is just some random effort to guess how test could look like. I think this setup is far from reality where you would build some realistic validation set. I can see the issue of having more \"badly\" labeled vs. \"well\" labeled data, but in this case having a specific validation dataset that is closer to test would produce way better solutions. ",
      "votes": 13
    },
    {
      "id": 3243849,
      "postDate": "2025-07-07T15:20:48.837Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> </p>\n<p>Sorry for opening this very old topic but I wanted to see your opinions on how this dataset is used in current research:</p>\n<p>In Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:<br>\n<a href=\"https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\" target=\"_blank\">https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf</a></p>\n<p>Here it seems that also report results from an internal validation set of 3,000 images in Table 1:<br>\n<a href=\"https://arxiv.org/pdf/2410.21969\" target=\"_blank\">https://arxiv.org/pdf/2410.21969</a></p>\n<p>It is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.</p>\n<p>I was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.</p>\n<p>What do you think? Do you ever manage to resolve the seeming misalignment between the Train and Test splits?</p>",
      "rawMarkdown": "Hi @gaborfodor and @philippsinger \n\nSorry for opening this very old topic but I wanted to see your opinions on how this dataset is used in current research:\n\nIn Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:\nhttps://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\n\nHere it seems that also report results from an internal validation set of 3,000 images in Table 1:\nhttps://arxiv.org/pdf/2410.21969\n\nIt is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.\n\nI was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.\n\nWhat do you think? Do you ever manage to resolve the seeming misalignment between the Train and Test splits?"
    }
  ],
  "comments": [
    {
      "id": 1257891,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-03-31T07:26:56.213000",
      "content": "<p>I fully agree with you <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> - there were some really weird things happening in this data.</p>\n<ul>\n<li><p>There was some extreme bias in what images which annotator annotates. I have the strong hypothesis that image distribution was not fully random. You could fit a classifier predicting the annotator for an image with very good accuracy. Actually we tried to remove the bias to better adjust to possible pattern changes in test, but it was basically impossible to remove the strong bias.</p></li>\n<li><p>The fact that each image either has three agreeing annotators for non-findings, or for findings, does not reflect any realistic use case. There should be images where some annotators say it is non-finding and others say it is a finding. There has to be some specific process in place.</p></li>\n<li><p>You have a strong difference between train and test labels, but you have no realistic chance of finding a really strong CV setup here. Public test is too small (300 images with large portion of nonfindings) to make too strong conclusions, and local CV is just some random effort to guess how test could look like. I think this setup is far from reality where you would build some realistic validation set. I can see the issue of having more \"badly\" labeled vs. \"well\" labeled data, but in this case having a specific validation dataset that is closer to test would produce way better solutions. </p></li>\n</ul>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 3243849,
      "author_name": "Joshua Bruton",
      "author_url": "",
      "post_date": "2025-07-07T15:20:48.837000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> and <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> </p>\n<p>Sorry for opening this very old topic but I wanted to see your opinions on how this dataset is used in current research:</p>\n<p>In Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:<br>\n<a href=\"https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\" target=\"_blank\">https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf</a></p>\n<p>Here it seems that also report results from an internal validation set of 3,000 images in Table 1:<br>\n<a href=\"https://arxiv.org/pdf/2410.21969\" target=\"_blank\">https://arxiv.org/pdf/2410.21969</a></p>\n<p>It is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.</p>\n<p>I was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.</p>\n<p>What do you think? Do you ever manage to resolve the seeming misalignment between the Train and Test splits?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1257745": "First, I would like to thank for the organizers for collecting the data and hosting this interesting challenge. It was my first kaggle competition with medical images and I learned a lot, special thanks for my teammates ( @pestipeti for object detection and @sandorkonya for radiology) for that!\nCongrats to the winners, I am really curious how they cracked this one and I am waiting for their summaries.\n\n\nMeanwhile I would like share a few serious bias/data leakage/structured noise issues we found that could hurt any real world application built on this dataset.  \n\n## Previous External Data Issues\nI started the competition by reviewing the organizers paper and the previous external chest X-ray datasets (CRX14, Chexpert) That way I found a few interesting blog posts about the possible problems with the already published datasets.\n\n* [Quick thoughts on ChestXray14, performance claims, and clinical tasks](https://lukeoakdenrayner.wordpress.com/2017/11/18/quick-thoughts-on-chestxray14-performance-claims-and-clinical-tasks/)\n* [Exploring the ChestXray14 dataset: problems](https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/)\n* [Half a million x-rays! First impressions of the Stanford and MIT chest x-ray datasets](https://lukeoakdenrayner.wordpress.com/2019/02/25/half-a-million-x-rays-first-impressions-of-the-stanford-and-mit-chest-x-ray-datasets/)\n\nThe organizers improved some of the already mentioned problems (NLP labels, image quality, AP/PA Xrays, multiple images per patient) but unfortunately they also introduced others.\n\n\n## DICOM meta data\n\nFirst we checked the DICOM meta data and extracted all the fields we could.\nWe trained a few XGB binary models for each class and surprisingly they gave **pretty good local validation score without using deep learning or even looking at the actual images!**\n\n![](https://storage.googleapis.com/vinbigdata/XGB.png)\n\nThat's weird! It means we have a serious leakage/bias in these fields.\nSome of the most important features were Age, Sex and **SourceApplicationEntityTitle (SAET)**. \nWhile Age and Sex could be important in medical diagnostics I can not explain the following normality rates by Sex:\n* F: 15%\n* M: 17%\n* O: 61%\n\n## No Finding class\nIt was discussed in the forums that the No Finding class had 100% agreement rate. That is very unusual in human annotation processes especially for such difficult and ambigous task.\nIt also means that the majority of the training set is not really annotated.\n\n## Annotator Group – SourceApplicationTitle\nAnalyzing the radiologist groups also lead to strange patterns R8-R9-R10  and R11-R17 always annotated together. R1-R7 always (except maybe one case) annotated No Finding.\nIf we check the annotations by SAET the pattern is a bit more clear\n![](https://storage.googleapis.com/vinbigdata/saet_rad.png)\n\n**The VITREA1 images did not have a single positive label as they were \"annotated\" by R1-R7.**\nUnfortunately these images took 53% of the test set. The vast majority of the positive training images came with Unknown SAET and were annotated by R8-R10, and unfortunately these images were 26% of the test set.\n\nI found the fact that many kagglers reported binary Normal/Abnormal models with AUC 0.992 very weird. We made a few tests and were able to learn the SAET almost perfectly. This kind of structured label bias could be very dangerous in real world medical application.\n\n## Class ambiguity\nThis one is the least serious one. I think at this point it is difficult to change too much on the pathology classes without loosing the ability to utilize the already collected hundreds of thousands public x-rays. It has been already discussed that the visual difference between classes could be quite difficult to define especially for labels consolidation / infiltration / atelectasis / pneumonia. \n\nJust another example, I wanted to understand better the ILD class, my quick google search resulted:\n\n*\"**Interstitial lung disease (ILD)** is another term for **pulmonary fibrosis**, which means “scarring” and “inflammation” of the interstitium (the tissue that surrounds the lung's air sacs, blood vessels and airways). This scarring makes the lung tissue stiff, which can make breathing difficult.Apr 26, 2018\"*\n\nIn a perfect world where annotators agree in 100% time each pathology would be found by all three annotators. Of course that is impossible even for easier annotation tasks (Chihuahua or muffin). On the training data the most obvious classes (Cardiomegaly, Aortic Enlargement) had perfect agreement 56-57% of the time while the most difficult classes (Consolidation, Other Lesion, Atelectasis) had only 10% triple agreement. That is quite a lot label noise to deal with and we have only 4K positive images in the training set.\n \n## Bounding box habits\n \nObject detection is much more complex than classification. It is also true for collecting ground truth labels.  Even if we decreased the minimum IOU to 0.4 lots of the annotated boxes did not agree at all. There were systematic differences between the radiologist groups too. R8-R10 usually used much smaller boxes and they were more consistent, while R11+ used less precise boxes. \n![](https://storage.googleapis.com/vinbigdata/AortaAnnotation.png)\n\nAnother systematic difference was that quite often R8-R10 only annotated the lower part of the heart while R11-R17 used larger boxes even bigger than the full heart.\n\n![](https://storage.googleapis.com/vinbigdata/Cardio.png)\nAs we have seen it is possible to learn these patterns with enough data but it does not make sense for real world applications.\n\n\n## Train-Test labeling differences\nI think using different train and test labeling method made the task even more difficult. We have already seen, that the the raw radiologist annotations had quite significant label noise. Using two unknown senior radiologists to clean the test set is not enough. Again, they could have different labeling habits (e.g. many precise small boxes, just one big box the lung section, lazy default box fusion etc.) and even keep/introduce more structured bias. \n\nI don't know how much would it cost to annotate the full training set as accurately as the test was labeled (My estimation is 1 min per image ~ 300 hours per radiologist) If that is too much, than providing 3K validation set annotated by the same procedure as the test set should be feasible. \n\nThese kind of train-test labeling differences lead to LB probing/LB overfitting and do not actually contribute to better models.\n\nThanks to all these issues, it is quite easy to train a simple baseline yolov5 model wtih 0.9+ AP local cross-validation for Aortic enlargement/Cardiomegaly then try it on the LB and realize it only achieves 0.15-0.4 AP. \n\nI have a few ideas that could explain some of the above issues (e.g. xrays coming from different hospitals/machines etc.) but I wonder how the host or other competitiors see it...\n",
    "1257891": "I fully agree with you @gaborfodor - there were some really weird things happening in this data.\n\n- There was some extreme bias in what images which annotator annotates. I have the strong hypothesis that image distribution was not fully random. You could fit a classifier predicting the annotator for an image with very good accuracy. Actually we tried to remove the bias to better adjust to possible pattern changes in test, but it was basically impossible to remove the strong bias.\n\n- The fact that each image either has three agreeing annotators for non-findings, or for findings, does not reflect any realistic use case. There should be images where some annotators say it is non-finding and others say it is a finding. There has to be some specific process in place.\n\n- You have a strong difference between train and test labels, but you have no realistic chance of finding a really strong CV setup here. Public test is too small (300 images with large portion of nonfindings) to make too strong conclusions, and local CV is just some random effort to guess how test could look like. I think this setup is far from reality where you would build some realistic validation set. I can see the issue of having more \"badly\" labeled vs. \"well\" labeled data, but in this case having a specific validation dataset that is closer to test would produce way better solutions. ",
    "3243849": "Hi @gaborfodor and @philippsinger \n\nSorry for opening this very old topic but I wanted to see your opinions on how this dataset is used in current research:\n\nIn Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:\nhttps://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\n\nHere it seems that also report results from an internal validation set of 3,000 images in Table 1:\nhttps://arxiv.org/pdf/2410.21969\n\nIt is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.\n\nI was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.\n\nWhat do you think? Do you ever manage to resolve the seeming misalignment between the Train and Test splits?"
  }
}