{
  "id": 207750,
  "title": "Labels Discussion / CheXNet Comparison",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/207750",
  "author_name": "Reuben Schmidt",
  "post_date": "2020-12-31T05:32:08.188000",
  "votes": 24,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Firstly, a big thank you to the organisers. For their <a href=\"https://storage.googleapis.com/kaggle-media/competitions/VinBigData/VinDr_CXR_data_paper.pdf\" target=\"_blank\">documentation</a>, sharing of the dataset as well as the <a href=\"https://github.com/vinbigdata-medical/vindr-cxr\" target=\"_blank\">scripts</a> used to create it.</p>\n<p>I wanted to highlight some similarity between this dataset and that of the <a href=\"https://arxiv.org/pdf/1711.05225.pdf\" target=\"_blank\">CheXNet paper</a>, which is one of the seminal papers in this field.</p>\n<p>Here are the labels from this competition compared to CheXNet's.</p>\n<table>\n<thead>\n<tr>\n<th>VinBigData</th>\n<th>CheXNet</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Aortic enlargement</td>\n<td>Hernia</td>\n</tr>\n<tr>\n<td><strong>Atelectasis</strong></td>\n<td><strong>Atelectasis</strong></td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>Pneumonia</td>\n</tr>\n<tr>\n<td><strong>Cardiomegaly</strong></td>\n<td><strong>Cardiomegaly</strong></td>\n</tr>\n<tr>\n<td><strong>Consolidation</strong></td>\n<td><strong>Consolidation</strong></td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>Edema</td>\n</tr>\n<tr>\n<td><strong>Infiltration</strong></td>\n<td><strong>Infiltration</strong></td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>Emphysema</td>\n</tr>\n<tr>\n<td>Nodule/Mass</td>\n<td>Nodule</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>Mass</td>\n</tr>\n<tr>\n<td><strong>Pleural effusion</strong></td>\n<td><strong>Effusion</strong></td>\n</tr>\n<tr>\n<td><strong>Pleural thickening</strong></td>\n<td><strong>Pleural Thickening</strong></td>\n</tr>\n<tr>\n<td><strong>Pneumothorax</strong></td>\n<td><strong>Pneumothorax</strong></td>\n</tr>\n<tr>\n<td><strong>Pulmonary fibrosis</strong></td>\n<td><strong>Fibrosis</strong></td>\n</tr>\n<tr>\n<td>No finding</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>CheXNet's labels have been <a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\" target=\"_blank\">fairly criticised</a> in the past, for consisting of some unusual choices for starters: the labels at times describe almost identical visual phenomena (like \"consolidation\" and \"infiltration\"). Other times the labels indicate clinical pathologies, rather than strictly findings on CXR (like \"pneumonia\" and \"emphysema\").</p>\n<p>This will probably be an interesting challenge for this dataset too. Perhaps in response to these criticisms of CheXNet / <a href=\"https://arxiv.org/pdf/1901.07031.pdf\" target=\"_blank\">CheXpert</a>, the organisers have done away with the labels \"pneumonia\" and \"emphysema\".</p>\n<p>Instead we have <a href=\"https://radiopaedia.org/articles/interstitial-lung-disease\" target=\"_blank\">ILD</a> (Interstitial Lung Disease), which is a diverse and varied set of findings and possible pathologies. We have calcification / nodule/mass / lung opacity / other lesion, which will have a bit of overlap I suspect.</p>\n<p>The good news is that these images have been specifically labelled by a group of 5 radiologists (rather than the labels being gathered from historic reports using NLP, like in the CheXNet paper), and that the test set labels are determined by consensus from all 5 radiologists. It will be interesting to see how people tackle this labelling approach.</p>",
  "messages": [
    {
      "id": 1133316,
      "postDate": "2020-12-31T05:32:08.187Z",
      "content": "<p>Firstly, a big thank you to the organisers. For their <a href=\"https://storage.googleapis.com/kaggle-media/competitions/VinBigData/VinDr_CXR_data_paper.pdf\" target=\"_blank\">documentation</a>, sharing of the dataset as well as the <a href=\"https://github.com/vinbigdata-medical/vindr-cxr\" target=\"_blank\">scripts</a> used to create it.</p>\n<p>I wanted to highlight some similarity between this dataset and that of the <a href=\"https://arxiv.org/pdf/1711.05225.pdf\" target=\"_blank\">CheXNet paper</a>, which is one of the seminal papers in this field.</p>\n<p>Here are the labels from this competition compared to CheXNet's.</p>\n<table>\n<thead>\n<tr>\n<th>VinBigData</th>\n<th>CheXNet</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Aortic enlargement</td>\n<td>Hernia</td>\n</tr>\n<tr>\n<td><strong>Atelectasis</strong></td>\n<td><strong>Atelectasis</strong></td>\n</tr>\n<tr>\n<td>Calcification</td>\n<td>Pneumonia</td>\n</tr>\n<tr>\n<td><strong>Cardiomegaly</strong></td>\n<td><strong>Cardiomegaly</strong></td>\n</tr>\n<tr>\n<td><strong>Consolidation</strong></td>\n<td><strong>Consolidation</strong></td>\n</tr>\n<tr>\n<td>ILD</td>\n<td>Edema</td>\n</tr>\n<tr>\n<td><strong>Infiltration</strong></td>\n<td><strong>Infiltration</strong></td>\n</tr>\n<tr>\n<td>Lung Opacity</td>\n<td>Emphysema</td>\n</tr>\n<tr>\n<td>Nodule/Mass</td>\n<td>Nodule</td>\n</tr>\n<tr>\n<td>Other lesion</td>\n<td>Mass</td>\n</tr>\n<tr>\n<td><strong>Pleural effusion</strong></td>\n<td><strong>Effusion</strong></td>\n</tr>\n<tr>\n<td><strong>Pleural thickening</strong></td>\n<td><strong>Pleural Thickening</strong></td>\n</tr>\n<tr>\n<td><strong>Pneumothorax</strong></td>\n<td><strong>Pneumothorax</strong></td>\n</tr>\n<tr>\n<td><strong>Pulmonary fibrosis</strong></td>\n<td><strong>Fibrosis</strong></td>\n</tr>\n<tr>\n<td>No finding</td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>CheXNet's labels have been <a href=\"https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/\" target=\"_blank\">fairly criticised</a> in the past, for consisting of some unusual choices for starters: the labels at times describe almost identical visual phenomena (like \"consolidation\" and \"infiltration\"). Other times the labels indicate clinical pathologies, rather than strictly findings on CXR (like \"pneumonia\" and \"emphysema\").</p>\n<p>This will probably be an interesting challenge for this dataset too. Perhaps in response to these criticisms of CheXNet / <a href=\"https://arxiv.org/pdf/1901.07031.pdf\" target=\"_blank\">CheXpert</a>, the organisers have done away with the labels \"pneumonia\" and \"emphysema\".</p>\n<p>Instead we have <a href=\"https://radiopaedia.org/articles/interstitial-lung-disease\" target=\"_blank\">ILD</a> (Interstitial Lung Disease), which is a diverse and varied set of findings and possible pathologies. We have calcification / nodule/mass / lung opacity / other lesion, which will have a bit of overlap I suspect.</p>\n<p>The good news is that these images have been specifically labelled by a group of 5 radiologists (rather than the labels being gathered from historic reports using NLP, like in the CheXNet paper), and that the test set labels are determined by consensus from all 5 radiologists. It will be interesting to see how people tackle this labelling approach.</p>",
      "rawMarkdown": "Firstly, a big thank you to the organisers. For their [documentation](https://storage.googleapis.com/kaggle-media/competitions/VinBigData/VinDr_CXR_data_paper.pdf), sharing of the dataset as well as the [scripts](https://github.com/vinbigdata-medical/vindr-cxr) used to create it.\n\nI wanted to highlight some similarity between this dataset and that of the [CheXNet paper](https://arxiv.org/pdf/1711.05225.pdf), which is one of the seminal papers in this field.\n\nHere are the labels from this competition compared to CheXNet's.\n\n| VinBigData | CheXNet |\n| --- | --- |\n|Aortic enlargement  | Hernia | \n| **Atelectasis**| **Atelectasis**| \n| Calcification  | Pneumonia | \n| **Cardiomegaly** | **Cardiomegaly** | \n| **Consolidation** | **Consolidation**| \n| ILD | Edema | \n| **Infiltration** | **Infiltration** | \n| Lung Opacity |  Emphysema | \n| Nodule/Mass  | Nodule | \n| Other lesion | Mass | \n| **Pleural effusion** | **Effusion** | \n| **Pleural thickening** | **Pleural Thickening** | \n| **Pneumothorax** | **Pneumothorax** | \n| **Pulmonary fibrosis**  |  **Fibrosis** |\n| No finding  |  |\n\nCheXNet's labels have been [fairly criticised](https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/) in the past, for consisting of some unusual choices for starters: the labels at times describe almost identical visual phenomena (like \"consolidation\" and \"infiltration\"). Other times the labels indicate clinical pathologies, rather than strictly findings on CXR (like \"pneumonia\" and \"emphysema\").\n\nThis will probably be an interesting challenge for this dataset too. Perhaps in response to these criticisms of CheXNet / [CheXpert](https://arxiv.org/pdf/1901.07031.pdf), the organisers have done away with the labels \"pneumonia\" and \"emphysema\".\n\nInstead we have [ILD](https://radiopaedia.org/articles/interstitial-lung-disease) (Interstitial Lung Disease), which is a diverse and varied set of findings and possible pathologies. We have calcification / nodule/mass / lung opacity / other lesion, which will have a bit of overlap I suspect.\n\nThe good news is that these images have been specifically labelled by a group of 5 radiologists (rather than the labels being gathered from historic reports using NLP, like in the CheXNet paper), and that the test set labels are determined by consensus from all 5 radiologists. It will be interesting to see how people tackle this labelling approach.",
      "votes": 24
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1133316": "Firstly, a big thank you to the organisers. For their [documentation](https://storage.googleapis.com/kaggle-media/competitions/VinBigData/VinDr_CXR_data_paper.pdf), sharing of the dataset as well as the [scripts](https://github.com/vinbigdata-medical/vindr-cxr) used to create it.\n\nI wanted to highlight some similarity between this dataset and that of the [CheXNet paper](https://arxiv.org/pdf/1711.05225.pdf), which is one of the seminal papers in this field.\n\nHere are the labels from this competition compared to CheXNet's.\n\n| VinBigData | CheXNet |\n| --- | --- |\n|Aortic enlargement  | Hernia | \n| **Atelectasis**| **Atelectasis**| \n| Calcification  | Pneumonia | \n| **Cardiomegaly** | **Cardiomegaly** | \n| **Consolidation** | **Consolidation**| \n| ILD | Edema | \n| **Infiltration** | **Infiltration** | \n| Lung Opacity |  Emphysema | \n| Nodule/Mass  | Nodule | \n| Other lesion | Mass | \n| **Pleural effusion** | **Effusion** | \n| **Pleural thickening** | **Pleural Thickening** | \n| **Pneumothorax** | **Pneumothorax** | \n| **Pulmonary fibrosis**  |  **Fibrosis** |\n| No finding  |  |\n\nCheXNet's labels have been [fairly criticised](https://lukeoakdenrayner.wordpress.com/2017/12/18/the-chestxray14-dataset-problems/) in the past, for consisting of some unusual choices for starters: the labels at times describe almost identical visual phenomena (like \"consolidation\" and \"infiltration\"). Other times the labels indicate clinical pathologies, rather than strictly findings on CXR (like \"pneumonia\" and \"emphysema\").\n\nThis will probably be an interesting challenge for this dataset too. Perhaps in response to these criticisms of CheXNet / [CheXpert](https://arxiv.org/pdf/1901.07031.pdf), the organisers have done away with the labels \"pneumonia\" and \"emphysema\".\n\nInstead we have [ILD](https://radiopaedia.org/articles/interstitial-lung-disease) (Interstitial Lung Disease), which is a diverse and varied set of findings and possible pathologies. We have calcification / nodule/mass / lung opacity / other lesion, which will have a bit of overlap I suspect.\n\nThe good news is that these images have been specifically labelled by a group of 5 radiologists (rather than the labels being gathered from historic reports using NLP, like in the CheXNet paper), and that the test set labels are determined by consensus from all 5 radiologists. It will be interesting to see how people tackle this labelling approach."
  }
}