{
  "id": 251250,
  "title": "Detailed analysis of the competition database - examination of consistency in annotation and data quality",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/251250",
  "author_name": "Weronika",
  "post_date": "2021-07-06T13:40:45.524000",
  "votes": 14,
  "comment_count": 2,
  "views": 0,
  "content": "<p>We would like to demonstrate how a series of simple tests for data imbalance exposes faults in the data acquisition and annotation process. We analyzed in detail data resources provided  to a Kaggle competition related to the detection of abnormalities in X-ray lung images. Complex models are able to learn artifacts and it is difficult to remove this bias during or after the training. Errors made at the data collection stage make it difficult to validate the model correctly.</p>\n<p>Problems in the training set provided for can be divided into 2 groups: problems related to consistency among radiologist and data quality problems.</p>\n<h1>Inconsistency among radiologists</h1>\n<h3>Unequal division of annotation work between radiologists</h3>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*2K8KEydxXMMsIS0xS937yg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*OOZe-c26E6OGG3Ae3S6YHw.png\">\n</p>\n<p>\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*ZeO39gXYABHKlV_nxNUnJA.png\">\n</p>\n<p><br>\n<em>Different label distributions among radiologists. The top left plot shows the number of images annotated by a radiologist grouped by whether the illness was found or not. The top right plot shows grouping by age and the bottom one grouping by&nbsp;sex.</em></p>\n<p>As visible in the figures above, the radiologists can be divided into three groups.</p>\n<p>The first group, R8-R10 worked on the same part of the X-ray dataset and annotated most of the images present in the dataset, both images with and without findings. Each radiologist annotated more than 6,000 images. Those three radiologists annotated 95% of all of the detected findings in this dataset.</p>\n<p>The next group R1-R7 did not detect almost any lesion (R2 found 3, the rest none).</p>\n<p>The last group, R11-R17. Each radiologist annotated less than 2,000 images with a high fraction of 'no findings' images.</p>\n<h3>Not clear annotation rules</h3>\n<p>Comparing the class labels given by different radiologists for a particular image, the consistency is remarkably low. In group R8-R10, radiologists (that annotated 95% of all findings) agreed with both colleagues on all classes only in 46% of images.</p>\n<table>\n<thead>\n<tr>\n<th>Radiologists</th>\n<th>R1-R7</th>\n<th>R8-R10</th>\n<th>R11-R17</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Agreed with at least one colleague on all classes</td>\n<td>100%</td>\n<td>69%</td>\n<td>96%</td>\n</tr>\n<tr>\n<td>Agreed with both colleagues on all classes</td>\n<td>100%</td>\n<td>46%</td>\n<td>94%</td>\n</tr>\n</tbody>\n</table>\n<h3>Different label for the same pathology</h3>\n<p>Another effect of unclear annotation rules is the significantly overlapping definitions of anomalies. A class ILD and Pulmonary fibrosis strongly overlap, similarly to Consolidation and Infiltration. The most vivid example is a \"lung opacity\", which covers six other classes!</p>\n<h3>Lesions present on chests with \"no findings\" label</h3>\n<p>Our expert radiologist analyzed 10 randomly selected images annotated by each of the seventeen radiologists (R1-R17). Surprisingly, we found out that although there was a~general consensus between dataset annotators when labeling \"no findings\", actually there are some anomalies that should be marked. He found some abnormalities in the images that were annotated as having 'no findings'. The exact numbers are visible in the table, and examples of mistakes are presented in the figure below.</p>\n<table>\n<thead>\n<tr>\n<th>radiologist's ID</th>\n<th>R1</th>\n<th>R2</th>\n<th>R3</th>\n<th>R4</th>\n<th>R5</th>\n<th>R6</th>\n<th>R7</th>\n<th>R8</th>\n<th>R9</th>\n<th>R10</th>\n<th>R11</th>\n<th>R12</th>\n<th>R13</th>\n<th>R14</th>\n<th>R15</th>\n<th>R16</th>\n<th>R17</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>number of errors</td>\n<td>0</td>\n<td>4</td>\n<td>0</td>\n<td>2</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n<td>1</td>\n<td>1</td>\n<td>2</td>\n<td>5</td>\n<td>5</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/1200/1*Rn2dc8KGUm5WF-y7CcNhIQ.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/1200/1*CHLlC_EbsFA41F_SRUYpsw.png\">\n</p>\n<p><em>Examples of lesions found on images checked by three radiologists and classified as No finding. The image on the left should be annotated as containing consolidation/pneumonia label, and the image on the right as Other lesion (actually dextrocardia).</em></p>\n<h3>One bounding box for all lesions of the same type, or one for each&nbsp;lesion</h3>\n<p>Some radiologists use a single box to cover few anomalies, others mark each anomaly separately.</p>\n<p>It influences model quality. The metric mAP at IoU 40, chosen for the competition, means that the predicted bounding box has to overlap with ground-truth box in at least 40%. The problem is that if radiologists' annotations (ground truth) do not meet this requirement, how is it possible to train an AI model with such noisy labels to get a good result.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*O5Mu5r0weAxzYtka34_UNg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*n-IPvCV-LD1P63ptdyhz3g.png\">\n</p>\n<p><em>Examples of inconsistency between radiologists related to the usage of a single box to mark many anomalies of the same class. On the left image, there are two big boxes each for the left and the right lung and many small boxes, on the right one, there is single box covering both&nbsp;lungs.</em></p>\n<h3>Different procedure of preparing train and test sets</h3>\n<p>The train and test sets were prepared differently. In both, the annotations were made independently by three radiologists for each image. According to (Nguyen, 2020), in the test set, there was an additional processing step. The labels were additionally verified and a consensus between two radiologists was reached.</p>\n<p>The problem is that there are considerable differences between radiologists. One approach is to select only critical findings and discard other annotations as unnecessary, which is acceptable for radiologists, but very challenging for nowadays ML model architectures. Typically, there is an assumption that a ML model should be trained on data similar to the target, and in order to deal with noise, more data is required.</p>\n<p>The second issue is the radiologist bias. From the training set analysis, we found out that most annotations were made by actually three radiologists (R8-R10). However, it is not known whether images annotated by those were used in the test dataset. This bias is reinforced by the additional two radiologists who made a consensus over annotations of three radiologists including standardization of label definitions. </p>\n<p>The role of two expert radiologists is unclear. It seems that those two only corrected annotations made by others. Their role should be much bigger, they are necessary to control if the annotation rules are well understood, and to clarify them if a new corner case arises. The standardized criteria for annotation should be prepared.</p>\n<h1>Data quality</h1>\n<h3>Lesions localization imbalance</h3>\n<p>In the database, there are 14 annotated anomalies. In regular clinician practice, all except 2 (aortic enlargement, cardiomegaly) are distributed similarly on both lungs sides. There should be a similar number of lesions in the right lung as well as in the left one. However, in the figure below, we placed heatmaps that should show the anomalies symmetrically appeared in both lungs.&nbsp;</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*pPVv6t-z0iHAsh6DEKpbqA.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*dsY21Wr3GaAx24xJfipJMQ.png\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*PWSKw4vbidpJCEG-rg3Ppg.png\">\n</p>\n<p><em>Examples of lesions that should be present symmetrically in both parts of the lungs. Before heatmaps were calculated, images from the training set were centered.</em></p>\n<h3>Children present in the&nbsp;dataset</h3>\n<p>In the training dataset, there are 107 images of children (ages 1–17). This might be a problem as child anatomy is different from adults (i.e., shape of heart, mediastinum, and bone structure) and so are the technical aspects of the child's X-ray (position of hands) (Hryniewska, 2020). The model might recognize such relationships. As children are not small adults, they should be removed in order not to introduce additional noise during model training.</p>\n<p>According to (Nguyen, 2020), pediatric X-rays should have been removed from the data during the data filtering step, but we found they were accidentally left.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*Y20ebmx-OecV3rDzRmI6Pg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\">\n</p>\n<p><em>Children's lungs​ versus adult's&nbsp;lungs</em></p>\n<h3>Two monochromatic color&nbsp;spaces</h3>\n<p>Another valid concern is Photometric Interpretation, which specifies the intended interpretation of the image pixel data. Some images are of type monochrome1 (17%) and some of monochrome2. The difference is that in the first case the lowest value of a pixel is interpreted as white and in the second case as black. This may produce some inefficient models when not taken into consideration.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*0w1en60KC3vhRDH2JSe9ng.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*lGZwA4bxbGegT9duFPpqLg.png\">\n</p>\n<p><em>monochrome1 versus monochrome2</em></p>\n<h3>Missing or wrong metadata in&nbsp;DICOMs</h3>\n<p>Some images suffer from missing data usually distributed in an extensive DICOM header, either from the lack of age or sex. Others don't have the correct data type (i.e., a letter instead of a number), so they were also classified by us as \"missing\". </p>\n<p>68% of the observations do not have information about age, and 17% about sex. The sex parameter is set to O (other) for 34% of the images. The rest of the dataset is fairly balanced (M: 26%, F: 23%).<br>\nThere is a lot of instances where the age is equal to 0 or is far greater than 100 (i.e. 238). This leaves us with only 25% of images with valid ages between 1-99.</p>\n<p>The lack of reliable information about age or sex is unfavorable because such attributes might be correlated with certain diseases, or having a disease at all. For example, for younger people, the probability of having lesions is significantly lower than for older people.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*ZhxqfhJn4ylEdVl3iuXGHg.png\">\n</p>\n<p><em>Density plots of age grouped by the existence of an illness. The probability of a young person having a lesion is lower than for the older&nbsp;person.</em></p>\n<h3>Parts of clothes present in the&nbsp;X-rays</h3>\n<p>Undesirable artifacts, presented in figures below, can be easily avoided during image acquisition, by asking the patient to remove all parts of the clothes that may influence X-ray imaging, for example, chains, bras, clothes with buttons, and zippers. If artifacts cannot be prevented, they can be removed during image preprocessing, before the image is shown to the model.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*jgKqYQU3cdjlewXM0YbUqw.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*T3dipvrbiUyu0dDI-h2UCQ.png\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*QD7p2ym39tjNHM1bQXMTQg.png\">\n</p>\n<p><em>Example of clothes artifacts. From the left, there are buttons, a zipper, a bone in a&nbsp;bra.</em></p>\n<h3>Letters present in the&nbsp;X-rays</h3>\n<p>Letters and/or annotations present in some lung images should be removed during preprocessing to prevent a neural network from learning those patterns. The model should learn how to differentiate labels by focusing on image features, not on descriptions in the images.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/800/1*sqf8rX2COsYb3yIx3DTmyg.png\">\n</p>\n<p><em>Example of letters artifacts.</em></p>\n<h1>Conclusions</h1>\n<p>The quality of a model is inherently bound to the quality of the data on which it is trained. Development of a reliable model should begin with data acquisition and annotation. At the model development stage, we cannot make the model fulfill all responsible AI and fairness rules if the data and their annotations are of insufficient quality.</p>\n<h1>Bibliography</h1>\n<p>Hryniewska, W., Bombiński, P., Szatkowski, P., Tomaszewska, P., Przelaskowski, A., &amp; Biecek, P. (2021). Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies. Pattern Recognition, 118, 108035. <a href=\"https://doi.org/10.1016/j.patcog.2021.108035\" target=\"_blank\">https://doi.org/10.1016/j.patcog.2021.108035</a></p>\n<p>Nguyen, H. Q., Lam, K., Le, L. T., Pham, H. H., Tran, D. Q., Nguyen, D. B., Le, D. D., Pham, C. M., Tong, H. T. T., Dinh, D. H., Do, C. D., Doan, L. T., Nguyen, C. N., Nguyen, B. T., Nguyen, Q. V., Hoang, A. D., Phan, H. N., Nguyen, A. T., Ho, P. H.,&nbsp;… Vu, V. (2020). VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations. <a href=\"http://arxiv.org/abs/2012.15029\" target=\"_blank\">http://arxiv.org/abs/2012.15029</a></p>",
  "messages": [
    {
      "id": 1378381,
      "postDate": "2021-07-06T13:40:45.523Z",
      "content": "<p>We would like to demonstrate how a series of simple tests for data imbalance exposes faults in the data acquisition and annotation process. We analyzed in detail data resources provided  to a Kaggle competition related to the detection of abnormalities in X-ray lung images. Complex models are able to learn artifacts and it is difficult to remove this bias during or after the training. Errors made at the data collection stage make it difficult to validate the model correctly.</p>\n<p>Problems in the training set provided for can be divided into 2 groups: problems related to consistency among radiologist and data quality problems.</p>\n<h1>Inconsistency among radiologists</h1>\n<h3>Unequal division of annotation work between radiologists</h3>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*2K8KEydxXMMsIS0xS937yg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*OOZe-c26E6OGG3Ae3S6YHw.png\">\n</p>\n<p>\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*ZeO39gXYABHKlV_nxNUnJA.png\">\n</p>\n<p><br>\n<em>Different label distributions among radiologists. The top left plot shows the number of images annotated by a radiologist grouped by whether the illness was found or not. The top right plot shows grouping by age and the bottom one grouping by&nbsp;sex.</em></p>\n<p>As visible in the figures above, the radiologists can be divided into three groups.</p>\n<p>The first group, R8-R10 worked on the same part of the X-ray dataset and annotated most of the images present in the dataset, both images with and without findings. Each radiologist annotated more than 6,000 images. Those three radiologists annotated 95% of all of the detected findings in this dataset.</p>\n<p>The next group R1-R7 did not detect almost any lesion (R2 found 3, the rest none).</p>\n<p>The last group, R11-R17. Each radiologist annotated less than 2,000 images with a high fraction of 'no findings' images.</p>\n<h3>Not clear annotation rules</h3>\n<p>Comparing the class labels given by different radiologists for a particular image, the consistency is remarkably low. In group R8-R10, radiologists (that annotated 95% of all findings) agreed with both colleagues on all classes only in 46% of images.</p>\n<table>\n<thead>\n<tr>\n<th>Radiologists</th>\n<th>R1-R7</th>\n<th>R8-R10</th>\n<th>R11-R17</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Agreed with at least one colleague on all classes</td>\n<td>100%</td>\n<td>69%</td>\n<td>96%</td>\n</tr>\n<tr>\n<td>Agreed with both colleagues on all classes</td>\n<td>100%</td>\n<td>46%</td>\n<td>94%</td>\n</tr>\n</tbody>\n</table>\n<h3>Different label for the same pathology</h3>\n<p>Another effect of unclear annotation rules is the significantly overlapping definitions of anomalies. A class ILD and Pulmonary fibrosis strongly overlap, similarly to Consolidation and Infiltration. The most vivid example is a \"lung opacity\", which covers six other classes!</p>\n<h3>Lesions present on chests with \"no findings\" label</h3>\n<p>Our expert radiologist analyzed 10 randomly selected images annotated by each of the seventeen radiologists (R1-R17). Surprisingly, we found out that although there was a~general consensus between dataset annotators when labeling \"no findings\", actually there are some anomalies that should be marked. He found some abnormalities in the images that were annotated as having 'no findings'. The exact numbers are visible in the table, and examples of mistakes are presented in the figure below.</p>\n<table>\n<thead>\n<tr>\n<th>radiologist's ID</th>\n<th>R1</th>\n<th>R2</th>\n<th>R3</th>\n<th>R4</th>\n<th>R5</th>\n<th>R6</th>\n<th>R7</th>\n<th>R8</th>\n<th>R9</th>\n<th>R10</th>\n<th>R11</th>\n<th>R12</th>\n<th>R13</th>\n<th>R14</th>\n<th>R15</th>\n<th>R16</th>\n<th>R17</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>number of errors</td>\n<td>0</td>\n<td>4</td>\n<td>0</td>\n<td>2</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n<td>1</td>\n<td>1</td>\n<td>2</td>\n<td>5</td>\n<td>5</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/1200/1*Rn2dc8KGUm5WF-y7CcNhIQ.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/1200/1*CHLlC_EbsFA41F_SRUYpsw.png\">\n</p>\n<p><em>Examples of lesions found on images checked by three radiologists and classified as No finding. The image on the left should be annotated as containing consolidation/pneumonia label, and the image on the right as Other lesion (actually dextrocardia).</em></p>\n<h3>One bounding box for all lesions of the same type, or one for each&nbsp;lesion</h3>\n<p>Some radiologists use a single box to cover few anomalies, others mark each anomaly separately.</p>\n<p>It influences model quality. The metric mAP at IoU 40, chosen for the competition, means that the predicted bounding box has to overlap with ground-truth box in at least 40%. The problem is that if radiologists' annotations (ground truth) do not meet this requirement, how is it possible to train an AI model with such noisy labels to get a good result.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*O5Mu5r0weAxzYtka34_UNg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*n-IPvCV-LD1P63ptdyhz3g.png\">\n</p>\n<p><em>Examples of inconsistency between radiologists related to the usage of a single box to mark many anomalies of the same class. On the left image, there are two big boxes each for the left and the right lung and many small boxes, on the right one, there is single box covering both&nbsp;lungs.</em></p>\n<h3>Different procedure of preparing train and test sets</h3>\n<p>The train and test sets were prepared differently. In both, the annotations were made independently by three radiologists for each image. According to (Nguyen, 2020), in the test set, there was an additional processing step. The labels were additionally verified and a consensus between two radiologists was reached.</p>\n<p>The problem is that there are considerable differences between radiologists. One approach is to select only critical findings and discard other annotations as unnecessary, which is acceptable for radiologists, but very challenging for nowadays ML model architectures. Typically, there is an assumption that a ML model should be trained on data similar to the target, and in order to deal with noise, more data is required.</p>\n<p>The second issue is the radiologist bias. From the training set analysis, we found out that most annotations were made by actually three radiologists (R8-R10). However, it is not known whether images annotated by those were used in the test dataset. This bias is reinforced by the additional two radiologists who made a consensus over annotations of three radiologists including standardization of label definitions. </p>\n<p>The role of two expert radiologists is unclear. It seems that those two only corrected annotations made by others. Their role should be much bigger, they are necessary to control if the annotation rules are well understood, and to clarify them if a new corner case arises. The standardized criteria for annotation should be prepared.</p>\n<h1>Data quality</h1>\n<h3>Lesions localization imbalance</h3>\n<p>In the database, there are 14 annotated anomalies. In regular clinician practice, all except 2 (aortic enlargement, cardiomegaly) are distributed similarly on both lungs sides. There should be a similar number of lesions in the right lung as well as in the left one. However, in the figure below, we placed heatmaps that should show the anomalies symmetrically appeared in both lungs.&nbsp;</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*pPVv6t-z0iHAsh6DEKpbqA.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*dsY21Wr3GaAx24xJfipJMQ.png\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*PWSKw4vbidpJCEG-rg3Ppg.png\">\n</p>\n<p><em>Examples of lesions that should be present symmetrically in both parts of the lungs. Before heatmaps were calculated, images from the training set were centered.</em></p>\n<h3>Children present in the&nbsp;dataset</h3>\n<p>In the training dataset, there are 107 images of children (ages 1–17). This might be a problem as child anatomy is different from adults (i.e., shape of heart, mediastinum, and bone structure) and so are the technical aspects of the child's X-ray (position of hands) (Hryniewska, 2020). The model might recognize such relationships. As children are not small adults, they should be removed in order not to introduce additional noise during model training.</p>\n<p>According to (Nguyen, 2020), pediatric X-rays should have been removed from the data during the data filtering step, but we found they were accidentally left.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*Y20ebmx-OecV3rDzRmI6Pg.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\">\n</p>\n<p><em>Children's lungs​ versus adult's&nbsp;lungs</em></p>\n<h3>Two monochromatic color&nbsp;spaces</h3>\n<p>Another valid concern is Photometric Interpretation, which specifies the intended interpretation of the image pixel data. Some images are of type monochrome1 (17%) and some of monochrome2. The difference is that in the first case the lowest value of a pixel is interpreted as white and in the second case as black. This may produce some inefficient models when not taken into consideration.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*0w1en60KC3vhRDH2JSe9ng.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*lGZwA4bxbGegT9duFPpqLg.png\">\n</p>\n<p><em>monochrome1 versus monochrome2</em></p>\n<h3>Missing or wrong metadata in&nbsp;DICOMs</h3>\n<p>Some images suffer from missing data usually distributed in an extensive DICOM header, either from the lack of age or sex. Others don't have the correct data type (i.e., a letter instead of a number), so they were also classified by us as \"missing\". </p>\n<p>68% of the observations do not have information about age, and 17% about sex. The sex parameter is set to O (other) for 34% of the images. The rest of the dataset is fairly balanced (M: 26%, F: 23%).<br>\nThere is a lot of instances where the age is equal to 0 or is far greater than 100 (i.e. 238). This leaves us with only 25% of images with valid ages between 1-99.</p>\n<p>The lack of reliable information about age or sex is unfavorable because such attributes might be correlated with certain diseases, or having a disease at all. For example, for younger people, the probability of having lesions is significantly lower than for older people.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*ZhxqfhJn4ylEdVl3iuXGHg.png\">\n</p>\n<p><em>Density plots of age grouped by the existence of an illness. The probability of a young person having a lesion is lower than for the older&nbsp;person.</em></p>\n<h3>Parts of clothes present in the&nbsp;X-rays</h3>\n<p>Undesirable artifacts, presented in figures below, can be easily avoided during image acquisition, by asking the patient to remove all parts of the clothes that may influence X-ray imaging, for example, chains, bras, clothes with buttons, and zippers. If artifacts cannot be prevented, they can be removed during image preprocessing, before the image is shown to the model.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*jgKqYQU3cdjlewXM0YbUqw.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*T3dipvrbiUyu0dDI-h2UCQ.png\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*QD7p2ym39tjNHM1bQXMTQg.png\">\n</p>\n<p><em>Example of clothes artifacts. From the left, there are buttons, a zipper, a bone in a&nbsp;bra.</em></p>\n<h3>Letters present in the&nbsp;X-rays</h3>\n<p>Letters and/or annotations present in some lung images should be removed during preprocessing to prevent a neural network from learning those patterns. The model should learn how to differentiate labels by focusing on image features, not on descriptions in the images.</p>\n<p>\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/800/1*sqf8rX2COsYb3yIx3DTmyg.png\">\n</p>\n<p><em>Example of letters artifacts.</em></p>\n<h1>Conclusions</h1>\n<p>The quality of a model is inherently bound to the quality of the data on which it is trained. Development of a reliable model should begin with data acquisition and annotation. At the model development stage, we cannot make the model fulfill all responsible AI and fairness rules if the data and their annotations are of insufficient quality.</p>\n<h1>Bibliography</h1>\n<p>Hryniewska, W., Bombiński, P., Szatkowski, P., Tomaszewska, P., Przelaskowski, A., &amp; Biecek, P. (2021). Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies. Pattern Recognition, 118, 108035. <a href=\"https://doi.org/10.1016/j.patcog.2021.108035\" target=\"_blank\">https://doi.org/10.1016/j.patcog.2021.108035</a></p>\n<p>Nguyen, H. Q., Lam, K., Le, L. T., Pham, H. H., Tran, D. Q., Nguyen, D. B., Le, D. D., Pham, C. M., Tong, H. T. T., Dinh, D. H., Do, C. D., Doan, L. T., Nguyen, C. N., Nguyen, B. T., Nguyen, Q. V., Hoang, A. D., Phan, H. N., Nguyen, A. T., Ho, P. H.,&nbsp;… Vu, V. (2020). VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations. <a href=\"http://arxiv.org/abs/2012.15029\" target=\"_blank\">http://arxiv.org/abs/2012.15029</a></p>",
      "rawMarkdown": "We would like to demonstrate how a series of simple tests for data imbalance exposes faults in the data acquisition and annotation process. We analyzed in detail data resources provided  to a Kaggle competition related to the detection of abnormalities in X-ray lung images. Complex models are able to learn artifacts and it is difficult to remove this bias during or after the training. Errors made at the data collection stage make it difficult to validate the model correctly.\n\nProblems in the training set provided for can be divided into 2 groups: problems related to consistency among radiologist and data quality problems.\n\n# Inconsistency among radiologists\n### Unequal division of annotation work between radiologists\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*2K8KEydxXMMsIS0xS937yg.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*OOZe-c26E6OGG3Ae3S6YHw.png\" width=\"45%\">\n</p>\n<p align=\"center\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*ZeO39gXYABHKlV_nxNUnJA.png\" width=\"60%\">\n</p>  \n*Different label distributions among radiologists. The top left plot shows the number of images annotated by a radiologist grouped by whether the illness was found or not. The top right plot shows grouping by age and the bottom one grouping by sex.*\n\nAs visible in the figures above, the radiologists can be divided into three groups.\n\nThe first group, R8-R10 worked on the same part of the X-ray dataset and annotated most of the images present in the dataset, both images with and without findings. Each radiologist annotated more than 6,000 images. Those three radiologists annotated 95% of all of the detected findings in this dataset.\n\nThe next group R1-R7 did not detect almost any lesion (R2 found 3, the rest none).\n\nThe last group, R11-R17. Each radiologist annotated less than 2,000 images with a high fraction of 'no findings' images.\n\n### Not clear annotation rules\nComparing the class labels given by different radiologists for a particular image, the consistency is remarkably low. In group R8-R10, radiologists (that annotated 95% of all findings) agreed with both colleagues on all classes only in 46% of images.\n\n| Radiologists |  R1-R7 | R8-R10 | R11-R17 |\n| --- | --- | --- | --- |\n| Agreed with at least one colleague on all classes | 100%  | 69%  | 96%  |\n| Agreed with both colleagues on all classes | 100%  | 46%  | 94%  |\n\n### Different label for the same pathology\nAnother effect of unclear annotation rules is the significantly overlapping definitions of anomalies. A class ILD and Pulmonary fibrosis strongly overlap, similarly to Consolidation and Infiltration. The most vivid example is a \"lung opacity\", which covers six other classes!\n\n### Lesions present on chests with \"no findings\" label\nOur expert radiologist analyzed 10 randomly selected images annotated by each of the seventeen radiologists (R1-R17). Surprisingly, we found out that although there was a~general consensus between dataset annotators when labeling \"no findings\", actually there are some anomalies that should be marked. He found some abnormalities in the images that were annotated as having 'no findings'. The exact numbers are visible in the table, and examples of mistakes are presented in the figure below.\n\n| radiologist's ID | R1 | R2 | R3 | R4 | R5 | R6 | R7 | R8 | R9 | R10 | R11 | R12 | R13 | R14 | R15 | R16 | R17 |   \n| --- | --- | --- | --- |\n| number of errors | 0 | 4 | 0 | 2 | 1 | 1 | 1 | 3 | 1 | 1 | 2 | 5 | 5 | 1 | 1 | 1 | 3 |\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/1200/1*Rn2dc8KGUm5WF-y7CcNhIQ.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/1200/1*CHLlC_EbsFA41F_SRUYpsw.png\" height=\"300\">\n</p>\n*Examples of lesions found on images checked by three radiologists and classified as No finding. The image on the left should be annotated as containing consolidation/pneumonia label, and the image on the right as Other lesion (actually dextrocardia).*\n\n### One bounding box for all lesions of the same type, or one for each lesion\nSome radiologists use a single box to cover few anomalies, others mark each anomaly separately.\n\nIt influences model quality. The metric mAP at IoU 40, chosen for the competition, means that the predicted bounding box has to overlap with ground-truth box in at least 40%. The problem is that if radiologists' annotations (ground truth) do not meet this requirement, how is it possible to train an AI model with such noisy labels to get a good result.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*O5Mu5r0weAxzYtka34_UNg.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*n-IPvCV-LD1P63ptdyhz3g.png\" width=\"45%\">\n</p>\n*Examples of inconsistency between radiologists related to the usage of a single box to mark many anomalies of the same class. On the left image, there are two big boxes each for the left and the right lung and many small boxes, on the right one, there is single box covering both lungs.*\n\n### Different procedure of preparing train and test sets\n\nThe train and test sets were prepared differently. In both, the annotations were made independently by three radiologists for each image. According to (Nguyen, 2020), in the test set, there was an additional processing step. The labels were additionally verified and a consensus between two radiologists was reached.\n\nThe problem is that there are considerable differences between radiologists. One approach is to select only critical findings and discard other annotations as unnecessary, which is acceptable for radiologists, but very challenging for nowadays ML model architectures. Typically, there is an assumption that a ML model should be trained on data similar to the target, and in order to deal with noise, more data is required.\n\nThe second issue is the radiologist bias. From the training set analysis, we found out that most annotations were made by actually three radiologists (R8-R10). However, it is not known whether images annotated by those were used in the test dataset. This bias is reinforced by the additional two radiologists who made a consensus over annotations of three radiologists including standardization of label definitions. \n\nThe role of two expert radiologists is unclear. It seems that those two only corrected annotations made by others. Their role should be much bigger, they are necessary to control if the annotation rules are well understood, and to clarify them if a new corner case arises. The standardized criteria for annotation should be prepared.\n\n# Data quality\n\n### Lesions localization imbalance\nIn the database, there are 14 annotated anomalies. In regular clinician practice, all except 2 (aortic enlargement, cardiomegaly) are distributed similarly on both lungs sides. There should be a similar number of lesions in the right lung as well as in the left one. However, in the figure below, we placed heatmaps that should show the anomalies symmetrically appeared in both lungs. \n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*pPVv6t-z0iHAsh6DEKpbqA.png\" width=\"30%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*dsY21Wr3GaAx24xJfipJMQ.png\" width=\"30%\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*PWSKw4vbidpJCEG-rg3Ppg.png\" width=\"30%\">\n</p>\n*Examples of lesions that should be present symmetrically in both parts of the lungs. Before heatmaps were calculated, images from the training set were centered.*\n\n### Children present in the dataset\nIn the training dataset, there are 107 images of children (ages 1–17). This might be a problem as child anatomy is different from adults (i.e., shape of heart, mediastinum, and bone structure) and so are the technical aspects of the child's X-ray (position of hands) (Hryniewska, 2020). The model might recognize such relationships. As children are not small adults, they should be removed in order not to introduce additional noise during model training.\n\nAccording to (Nguyen, 2020), pediatric X-rays should have been removed from the data during the data filtering step, but we found they were accidentally left.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*Y20ebmx-OecV3rDzRmI6Pg.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\" height=\"300\">\n</p>\n*Children's lungs​ versus adult's lungs*\n\n### Two monochromatic color spaces\nAnother valid concern is Photometric Interpretation, which specifies the intended interpretation of the image pixel data. Some images are of type monochrome1 (17%) and some of monochrome2. The difference is that in the first case the lowest value of a pixel is interpreted as white and in the second case as black. This may produce some inefficient models when not taken into consideration.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*0w1en60KC3vhRDH2JSe9ng.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*lGZwA4bxbGegT9duFPpqLg.png\" width=\"45%\">\n</p>\n*monochrome1 versus monochrome2*\n\n### Missing or wrong metadata in DICOMs\nSome images suffer from missing data usually distributed in an extensive DICOM header, either from the lack of age or sex. Others don't have the correct data type (i.e., a letter instead of a number), so they were also classified by us as \"missing\". \n\n68% of the observations do not have information about age, and 17% about sex. The sex parameter is set to O (other) for 34% of the images. The rest of the dataset is fairly balanced (M: 26%, F: 23%).\nThere is a lot of instances where the age is equal to 0 or is far greater than 100 (i.e. 238). This leaves us with only 25% of images with valid ages between 1-99.\n\nThe lack of reliable information about age or sex is unfavorable because such attributes might be correlated with certain diseases, or having a disease at all. For example, for younger people, the probability of having lesions is significantly lower than for older people.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*ZhxqfhJn4ylEdVl3iuXGHg.png\" width=\"100%\">\n</p>\n*Density plots of age grouped by the existence of an illness. The probability of a young person having a lesion is lower than for the older person.*\n\n### Parts of clothes present in the X-rays\nUndesirable artifacts, presented in figures below, can be easily avoided during image acquisition, by asking the patient to remove all parts of the clothes that may influence X-ray imaging, for example, chains, bras, clothes with buttons, and zippers. If artifacts cannot be prevented, they can be removed during image preprocessing, before the image is shown to the model.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*jgKqYQU3cdjlewXM0YbUqw.png\" width=\"30%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*T3dipvrbiUyu0dDI-h2UCQ.png\" width=\"30%\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*QD7p2ym39tjNHM1bQXMTQg.png\" width=\"30%\">\n</p>\n*Example of clothes artifacts. From the left, there are buttons, a zipper, a bone in a bra.*\n\n### Letters present in the X-rays\nLetters and/or annotations present in some lung images should be removed during preprocessing to prevent a neural network from learning those patterns. The model should learn how to differentiate labels by focusing on image features, not on descriptions in the images.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/800/1*sqf8rX2COsYb3yIx3DTmyg.png\" height=\"300\">\n</p>\n*Example of letters artifacts.*\n\n# Conclusions\nThe quality of a model is inherently bound to the quality of the data on which it is trained. Development of a reliable model should begin with data acquisition and annotation. At the model development stage, we cannot make the model fulfill all responsible AI and fairness rules if the data and their annotations are of insufficient quality.\n\n# Bibliography\nHryniewska, W., Bombiński, P., Szatkowski, P., Tomaszewska, P., Przelaskowski, A., & Biecek, P. (2021). Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies. Pattern Recognition, 118, 108035. https://doi.org/10.1016/j.patcog.2021.108035\n\nNguyen, H. Q., Lam, K., Le, L. T., Pham, H. H., Tran, D. Q., Nguyen, D. B., Le, D. D., Pham, C. M., Tong, H. T. T., Dinh, D. H., Do, C. D., Doan, L. T., Nguyen, C. N., Nguyen, B. T., Nguyen, Q. V., Hoang, A. D., Phan, H. N., Nguyen, A. T., Ho, P. H., … Vu, V. (2020). VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations. http://arxiv.org/abs/2012.15029",
      "votes": 14
    },
    {
      "id": 1379222,
      "postDate": "2021-07-07T07:19:13.343Z",
      "content": "<p>Thanks for sharing, this is great information!</p>",
      "rawMarkdown": "Thanks for sharing, this is great information!"
    },
    {
      "id": 1432913,
      "postDate": "2021-08-03T13:37:25.200Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1379222,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-07-07T07:19:13.343000",
      "content": "<p>Thanks for sharing, this is great information!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1432913,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-03T13:37:25.200000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1378381": "We would like to demonstrate how a series of simple tests for data imbalance exposes faults in the data acquisition and annotation process. We analyzed in detail data resources provided  to a Kaggle competition related to the detection of abnormalities in X-ray lung images. Complex models are able to learn artifacts and it is difficult to remove this bias during or after the training. Errors made at the data collection stage make it difficult to validate the model correctly.\n\nProblems in the training set provided for can be divided into 2 groups: problems related to consistency among radiologist and data quality problems.\n\n# Inconsistency among radiologists\n### Unequal division of annotation work between radiologists\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*2K8KEydxXMMsIS0xS937yg.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*OOZe-c26E6OGG3Ae3S6YHw.png\" width=\"45%\">\n</p>\n<p align=\"center\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*ZeO39gXYABHKlV_nxNUnJA.png\" width=\"60%\">\n</p>  \n*Different label distributions among radiologists. The top left plot shows the number of images annotated by a radiologist grouped by whether the illness was found or not. The top right plot shows grouping by age and the bottom one grouping by sex.*\n\nAs visible in the figures above, the radiologists can be divided into three groups.\n\nThe first group, R8-R10 worked on the same part of the X-ray dataset and annotated most of the images present in the dataset, both images with and without findings. Each radiologist annotated more than 6,000 images. Those three radiologists annotated 95% of all of the detected findings in this dataset.\n\nThe next group R1-R7 did not detect almost any lesion (R2 found 3, the rest none).\n\nThe last group, R11-R17. Each radiologist annotated less than 2,000 images with a high fraction of 'no findings' images.\n\n### Not clear annotation rules\nComparing the class labels given by different radiologists for a particular image, the consistency is remarkably low. In group R8-R10, radiologists (that annotated 95% of all findings) agreed with both colleagues on all classes only in 46% of images.\n\n| Radiologists |  R1-R7 | R8-R10 | R11-R17 |\n| --- | --- | --- | --- |\n| Agreed with at least one colleague on all classes | 100%  | 69%  | 96%  |\n| Agreed with both colleagues on all classes | 100%  | 46%  | 94%  |\n\n### Different label for the same pathology\nAnother effect of unclear annotation rules is the significantly overlapping definitions of anomalies. A class ILD and Pulmonary fibrosis strongly overlap, similarly to Consolidation and Infiltration. The most vivid example is a \"lung opacity\", which covers six other classes!\n\n### Lesions present on chests with \"no findings\" label\nOur expert radiologist analyzed 10 randomly selected images annotated by each of the seventeen radiologists (R1-R17). Surprisingly, we found out that although there was a~general consensus between dataset annotators when labeling \"no findings\", actually there are some anomalies that should be marked. He found some abnormalities in the images that were annotated as having 'no findings'. The exact numbers are visible in the table, and examples of mistakes are presented in the figure below.\n\n| radiologist's ID | R1 | R2 | R3 | R4 | R5 | R6 | R7 | R8 | R9 | R10 | R11 | R12 | R13 | R14 | R15 | R16 | R17 |   \n| --- | --- | --- | --- |\n| number of errors | 0 | 4 | 0 | 2 | 1 | 1 | 1 | 3 | 1 | 1 | 2 | 5 | 5 | 1 | 1 | 1 | 3 |\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/1200/1*Rn2dc8KGUm5WF-y7CcNhIQ.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/1200/1*CHLlC_EbsFA41F_SRUYpsw.png\" height=\"300\">\n</p>\n*Examples of lesions found on images checked by three radiologists and classified as No finding. The image on the left should be annotated as containing consolidation/pneumonia label, and the image on the right as Other lesion (actually dextrocardia).*\n\n### One bounding box for all lesions of the same type, or one for each lesion\nSome radiologists use a single box to cover few anomalies, others mark each anomaly separately.\n\nIt influences model quality. The metric mAP at IoU 40, chosen for the competition, means that the predicted bounding box has to overlap with ground-truth box in at least 40%. The problem is that if radiologists' annotations (ground truth) do not meet this requirement, how is it possible to train an AI model with such noisy labels to get a good result.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*O5Mu5r0weAxzYtka34_UNg.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*n-IPvCV-LD1P63ptdyhz3g.png\" width=\"45%\">\n</p>\n*Examples of inconsistency between radiologists related to the usage of a single box to mark many anomalies of the same class. On the left image, there are two big boxes each for the left and the right lung and many small boxes, on the right one, there is single box covering both lungs.*\n\n### Different procedure of preparing train and test sets\n\nThe train and test sets were prepared differently. In both, the annotations were made independently by three radiologists for each image. According to (Nguyen, 2020), in the test set, there was an additional processing step. The labels were additionally verified and a consensus between two radiologists was reached.\n\nThe problem is that there are considerable differences between radiologists. One approach is to select only critical findings and discard other annotations as unnecessary, which is acceptable for radiologists, but very challenging for nowadays ML model architectures. Typically, there is an assumption that a ML model should be trained on data similar to the target, and in order to deal with noise, more data is required.\n\nThe second issue is the radiologist bias. From the training set analysis, we found out that most annotations were made by actually three radiologists (R8-R10). However, it is not known whether images annotated by those were used in the test dataset. This bias is reinforced by the additional two radiologists who made a consensus over annotations of three radiologists including standardization of label definitions. \n\nThe role of two expert radiologists is unclear. It seems that those two only corrected annotations made by others. Their role should be much bigger, they are necessary to control if the annotation rules are well understood, and to clarify them if a new corner case arises. The standardized criteria for annotation should be prepared.\n\n# Data quality\n\n### Lesions localization imbalance\nIn the database, there are 14 annotated anomalies. In regular clinician practice, all except 2 (aortic enlargement, cardiomegaly) are distributed similarly on both lungs sides. There should be a similar number of lesions in the right lung as well as in the left one. However, in the figure below, we placed heatmaps that should show the anomalies symmetrically appeared in both lungs. \n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/400/1*pPVv6t-z0iHAsh6DEKpbqA.png\" width=\"30%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*dsY21Wr3GaAx24xJfipJMQ.png\" width=\"30%\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*PWSKw4vbidpJCEG-rg3Ppg.png\" width=\"30%\">\n</p>\n*Examples of lesions that should be present symmetrically in both parts of the lungs. Before heatmaps were calculated, images from the training set were centered.*\n\n### Children present in the dataset\nIn the training dataset, there are 107 images of children (ages 1–17). This might be a problem as child anatomy is different from adults (i.e., shape of heart, mediastinum, and bone structure) and so are the technical aspects of the child's X-ray (position of hands) (Hryniewska, 2020). The model might recognize such relationships. As children are not small adults, they should be removed in order not to introduce additional noise during model training.\n\nAccording to (Nguyen, 2020), pediatric X-rays should have been removed from the data during the data filtering step, but we found they were accidentally left.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*Y20ebmx-OecV3rDzRmI6Pg.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\" height=\"300\">\n</p>\n*Children's lungs​ versus adult's lungs*\n\n### Two monochromatic color spaces\nAnother valid concern is Photometric Interpretation, which specifies the intended interpretation of the image pixel data. Some images are of type monochrome1 (17%) and some of monochrome2. The difference is that in the first case the lowest value of a pixel is interpreted as white and in the second case as black. This may produce some inefficient models when not taken into consideration.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*0w1en60KC3vhRDH2JSe9ng.png\" width=\"45%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*lGZwA4bxbGegT9duFPpqLg.png\" width=\"45%\">\n</p>\n*monochrome1 versus monochrome2*\n\n### Missing or wrong metadata in DICOMs\nSome images suffer from missing data usually distributed in an extensive DICOM header, either from the lack of age or sex. Others don't have the correct data type (i.e., a letter instead of a number), so they were also classified by us as \"missing\". \n\n68% of the observations do not have information about age, and 17% about sex. The sex parameter is set to O (other) for 34% of the images. The rest of the dataset is fairly balanced (M: 26%, F: 23%).\nThere is a lot of instances where the age is equal to 0 or is far greater than 100 (i.e. 238). This leaves us with only 25% of images with valid ages between 1-99.\n\nThe lack of reliable information about age or sex is unfavorable because such attributes might be correlated with certain diseases, or having a disease at all. For example, for younger people, the probability of having lesions is significantly lower than for older people.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/800/1*ZhxqfhJn4ylEdVl3iuXGHg.png\" width=\"100%\">\n</p>\n*Density plots of age grouped by the existence of an illness. The probability of a young person having a lesion is lower than for the older person.*\n\n### Parts of clothes present in the X-rays\nUndesirable artifacts, presented in figures below, can be easily avoided during image acquisition, by asking the patient to remove all parts of the clothes that may influence X-ray imaging, for example, chains, bras, clothes with buttons, and zippers. If artifacts cannot be prevented, they can be removed during image preprocessing, before the image is shown to the model.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*jgKqYQU3cdjlewXM0YbUqw.png\" width=\"30%\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/600/1*T3dipvrbiUyu0dDI-h2UCQ.png\" width=\"30%\">\n<img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/400/1*QD7p2ym39tjNHM1bQXMTQg.png\" width=\"30%\">\n</p>\n*Example of clothes artifacts. From the left, there are buttons, a zipper, a bone in a bra.*\n\n### Letters present in the X-rays\nLetters and/or annotations present in some lung images should be removed during preprocessing to prevent a neural network from learning those patterns. The model should learn how to differentiate labels by focusing on image features, not on descriptions in the images.\n\n<p align=\"center\">\n  <img alt=\"Light\" src=\"https://cdn-images-1.medium.com/max/600/1*3yD9hdBUf9LoLXiM7ye8ow.png\" height=\"300\">\n&nbsp; &nbsp; &nbsp; &nbsp;\n  <img alt=\"Dark\" src=\"https://cdn-images-1.medium.com/max/800/1*sqf8rX2COsYb3yIx3DTmyg.png\" height=\"300\">\n</p>\n*Example of letters artifacts.*\n\n# Conclusions\nThe quality of a model is inherently bound to the quality of the data on which it is trained. Development of a reliable model should begin with data acquisition and annotation. At the model development stage, we cannot make the model fulfill all responsible AI and fairness rules if the data and their annotations are of insufficient quality.\n\n# Bibliography\nHryniewska, W., Bombiński, P., Szatkowski, P., Tomaszewska, P., Przelaskowski, A., & Biecek, P. (2021). Checklist for responsible deep learning modeling of medical images based on COVID-19 detection studies. Pattern Recognition, 118, 108035. https://doi.org/10.1016/j.patcog.2021.108035\n\nNguyen, H. Q., Lam, K., Le, L. T., Pham, H. H., Tran, D. Q., Nguyen, D. B., Le, D. D., Pham, C. M., Tong, H. T. T., Dinh, D. H., Do, C. D., Doan, L. T., Nguyen, C. N., Nguyen, B. T., Nguyen, Q. V., Hoang, A. D., Phan, H. N., Nguyen, A. T., Ho, P. H., … Vu, V. (2020). VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations. http://arxiv.org/abs/2012.15029",
    "1379222": "Thanks for sharing, this is great information!",
    "1432913": ""
  }
}