{
  "id": 216181,
  "title": "Some radiologists only ever reviewed images with \"no findings\"?!",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/216181",
  "author_name": "Björn",
  "post_date": "2021-02-02T00:02:39.448000",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/215444\" target=\"_blank\">other discussion</a> prompted me to look at the performance of the 17 different radiologists. I found something interesting (<a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">link to notebook</a>): To my surprise several of them had a 100% match in terms of <code>class_id</code> values assigned to each image with the other three radiologists looking a the same image.</p>\n<p>It turns out that radiologists that 100% agree with other radiologists only ever assigned <code>class_id = 14</code> (=\"no finding\"). I'd guess X-rays that were a-priori identified (how?) to be without a finding and were sent to a subset of radiologists, especially <code>R1</code> to <code>R7</code> (and sometimes others, but , <code>R8</code>, <code>R9</code> and <code>R10</code> mostly did not get those images).</p>\n<p>None of this is mentioned in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">the paper</a>. So I'm wondering whether there's something I overlooked and whether there's other things I did miss about the training process. </p>\n<p>It also makes me think how much we should trust the \"No finding\" images… If someone has to go through 1500 to 3000 images, one after another and there's just never anything there, do they get tired and sloppy (and their colleagues, too)? I would have thought that mixing things up would be a good idea to make sure the radiologists are kept engaged/interested/\"on their toes\" (or however else you want to describe it).</p>\n<table>\n<thead>\n<tr>\n<th>rad_id</th>\n<th>Images with no finding</th>\n<th>Percent with no finding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>R1</td>\n<td>1995</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R2</td>\n<td>3118</td>\n<td>99.9%</td>\n</tr>\n<tr>\n<td>R3</td>\n<td>2285</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R4</td>\n<td>1513</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R5</td>\n<td>2783</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R6</td>\n<td>2041</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R7</td>\n<td>1733</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R8</td>\n<td>2436</td>\n<td>19.97%</td>\n</tr>\n<tr>\n<td>R9</td>\n<td>1979</td>\n<td>12.6%</td>\n</tr>\n<tr>\n<td>R10</td>\n<td>2321</td>\n<td>17.46%</td>\n</tr>\n<tr>\n<td>R11</td>\n<td>1413</td>\n<td>84.61%</td>\n</tr>\n<tr>\n<td>R12</td>\n<td>1580</td>\n<td>91.38%</td>\n</tr>\n<tr>\n<td>R13</td>\n<td>1505</td>\n<td>82.51%</td>\n</tr>\n<tr>\n<td>R14</td>\n<td>1300</td>\n<td>80.05%</td>\n</tr>\n<tr>\n<td>R15</td>\n<td>1508</td>\n<td>82.72%</td>\n</tr>\n<tr>\n<td>R16</td>\n<td>1565</td>\n<td>88.77%</td>\n</tr>\n<tr>\n<td>R17</td>\n<td>743</td>\n<td>91.5%</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 1181501,
      "postDate": "2021-02-02T00:02:39.450Z",
      "content": "<p>This <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/215444\" target=\"_blank\">other discussion</a> prompted me to look at the performance of the 17 different radiologists. I found something interesting (<a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">link to notebook</a>): To my surprise several of them had a 100% match in terms of <code>class_id</code> values assigned to each image with the other three radiologists looking a the same image.</p>\n<p>It turns out that radiologists that 100% agree with other radiologists only ever assigned <code>class_id = 14</code> (=\"no finding\"). I'd guess X-rays that were a-priori identified (how?) to be without a finding and were sent to a subset of radiologists, especially <code>R1</code> to <code>R7</code> (and sometimes others, but , <code>R8</code>, <code>R9</code> and <code>R10</code> mostly did not get those images).</p>\n<p>None of this is mentioned in <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">the paper</a>. So I'm wondering whether there's something I overlooked and whether there's other things I did miss about the training process. </p>\n<p>It also makes me think how much we should trust the \"No finding\" images… If someone has to go through 1500 to 3000 images, one after another and there's just never anything there, do they get tired and sloppy (and their colleagues, too)? I would have thought that mixing things up would be a good idea to make sure the radiologists are kept engaged/interested/\"on their toes\" (or however else you want to describe it).</p>\n<table>\n<thead>\n<tr>\n<th>rad_id</th>\n<th>Images with no finding</th>\n<th>Percent with no finding</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>R1</td>\n<td>1995</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R2</td>\n<td>3118</td>\n<td>99.9%</td>\n</tr>\n<tr>\n<td>R3</td>\n<td>2285</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R4</td>\n<td>1513</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R5</td>\n<td>2783</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R6</td>\n<td>2041</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R7</td>\n<td>1733</td>\n<td>100%</td>\n</tr>\n<tr>\n<td>R8</td>\n<td>2436</td>\n<td>19.97%</td>\n</tr>\n<tr>\n<td>R9</td>\n<td>1979</td>\n<td>12.6%</td>\n</tr>\n<tr>\n<td>R10</td>\n<td>2321</td>\n<td>17.46%</td>\n</tr>\n<tr>\n<td>R11</td>\n<td>1413</td>\n<td>84.61%</td>\n</tr>\n<tr>\n<td>R12</td>\n<td>1580</td>\n<td>91.38%</td>\n</tr>\n<tr>\n<td>R13</td>\n<td>1505</td>\n<td>82.51%</td>\n</tr>\n<tr>\n<td>R14</td>\n<td>1300</td>\n<td>80.05%</td>\n</tr>\n<tr>\n<td>R15</td>\n<td>1508</td>\n<td>82.72%</td>\n</tr>\n<tr>\n<td>R16</td>\n<td>1565</td>\n<td>88.77%</td>\n</tr>\n<tr>\n<td>R17</td>\n<td>743</td>\n<td>91.5%</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "This [other discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/215444) prompted me to look at the performance of the 17 different radiologists. I found something interesting ([link to notebook](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/#1.-Labelling-process-that-created-the-training-and-test-data)): To my surprise several of them had a 100% match in terms of `class_id` values assigned to each image with the other three radiologists looking a the same image.\n\nIt turns out that radiologists that 100% agree with other radiologists only ever assigned `class_id = 14` (=\"no finding\"). I'd guess X-rays that were a-priori identified (how?) to be without a finding and were sent to a subset of radiologists, especially `R1` to `R7` (and sometimes others, but , `R8`, `R9` and `R10` mostly did not get those images).\n\nNone of this is mentioned in [the paper](https://arxiv.org/pdf/2012.15029.pdf). So I'm wondering whether there's something I overlooked and whether there's other things I did miss about the training process. \n\nIt also makes me think how much we should trust the \"No finding\" images... If someone has to go through 1500 to 3000 images, one after another and there's just never anything there, do they get tired and sloppy (and their colleagues, too)? I would have thought that mixing things up would be a good idea to make sure the radiologists are kept engaged/interested/\"on their toes\" (or however else you want to describe it).\n\n| rad_id | Images with no finding | Percent with no finding |\n| --- | --- |\n|R1 |\t1995 |100% |\n|R2\t| 3118 | 99.9% |\n|R3\t| 2285 | 100% |\n|R4\t| 1513 | 100% |\n|R5\t| 2783 | 100% |\n|R6\t| 2041 | 100% |\n|R7\t| 1733 | 100% |\n|R8\t| 2436 | 19.97% |\n|R9\t| 1979 | 12.6% |\n|R10\t| 2321 | 17.46% |\n|R11 | 1413 | 84.61% |\n|R12|1580 | 91.38% |\n|R13\t|1505 | 82.51% |\n|R14\t|1300 | 80.05% |\n|R15\t|1508 | 82.72% |\n|R16\t|1565 | 88.77% |\n|R17 | 743 | 91.5% |\n",
      "votes": 10
    },
    {
      "id": 1211830,
      "postDate": "2021-02-20T15:39:43.717Z",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>,<br>\nthe normal daily routine is different in every hospital, there are houses where the junior radiologists send each and every report to second reading and places where only the \"questionable\" ones are sent to second reader.</p>\n<p>I think that the \"normal\" tagged images come from a screening ( young ppl must have a chest x-ray before starting work to exclude pulmonary TB ) and these are with high % without finding.</p>",
      "rawMarkdown": "@bjoernholzhauer,\nthe normal daily routine is different in every hospital, there are houses where the junior radiologists send each and every report to second reading and places where only the \"questionable\" ones are sent to second reader.\n\nI think that the \"normal\" tagged images come from a screening ( young ppl must have a chest x-ray before starting work to exclude pulmonary TB ) and these are with high % without finding.",
      "votes": 1,
      "replies": [
        {
          "id": 1211861,
          "postDate": "2021-02-20T16:09:58.047Z",
          "content": "<p>Thanks, that sounds plausible. It was not clearly described in the paper on the dataset. Largely a good job on those, given that no findings seems to then have mostly been confirmed. But, probably raises the risk that a subtle finding was missed originally and that the reviewers for this dataset got some fatigue after clicking though one\"no finding\" after another and missed it, too, I'd guess?</p>",
          "rawMarkdown": "Thanks, that sounds plausible. It was not clearly described in the paper on the dataset. Largely a good job on those, given that no findings seems to then have mostly been confirmed. But, probably raises the risk that a subtle finding was missed originally and that the reviewers for this dataset got some fatigue after clicking though one\"no finding\" after another and missed it, too, I'd guess?"
        },
        {
          "id": 1211891,
          "postDate": "2021-02-20T16:29:04.213Z",
          "content": "<p>It is also just guessing from my side… </p>\n<p>Sure, it is always a possibility to have subtle finding on them though!</p>",
          "rawMarkdown": "It is also just guessing from my side... \n\nSure, it is always a possibility to have subtle finding on them though!\n"
        }
      ]
    },
    {
      "id": 1184022,
      "postDate": "2021-02-03T11:05:01.830Z",
      "content": "<p>Interesting!</p>\n<p>Disregarding the &gt;99% no finding radiologists, it seems unlikely that R8, R9, R10 get no findings 80% of the time.</p>\n<p>Unless the xrays were distributed non-randomly.</p>",
      "rawMarkdown": "Interesting!\n\nDisregarding the >99% no finding radiologists, it seems unlikely that R8, R9, R10 get no findings <20% of the time while R11-R17 get no findings >80% of the time.\n\nUnless the xrays were distributed non-randomly.",
      "votes": 1,
      "replies": [
        {
          "id": 1184065,
          "postDate": "2021-02-03T11:24:54.730Z",
          "content": "<p>Yes, I think it's pretty clear that the x-rays were assigned to radiologists in a non-random way. I'm guessing that they split by the original assigned diagnosis at the two hospitals where they got it from (or perhaps they ran some existing algorithm over it) and then somehow systematically assigned \"no findings\" x-rays mostly to one set of reviewers and ones with findings to others (or something like that).</p>",
          "rawMarkdown": "Yes, I think it's pretty clear that the x-rays were assigned to radiologists in a non-random way. I'm guessing that they split by the original assigned diagnosis at the two hospitals where they got it from (or perhaps they ran some existing algorithm over it) and then somehow systematically assigned \"no findings\" x-rays mostly to one set of reviewers and ones with findings to others (or something like that).",
          "votes": 2
        },
        {
          "id": 1184083,
          "postDate": "2021-02-03T11:33:09.537Z",
          "content": "<p>It's odd that from what I can see, all radiologists agree when there's \"no finding.\"</p>",
          "rawMarkdown": "It's odd that from what I can see, all radiologists agree when there's \"no finding.\"",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1211830,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2021-02-20T15:39:43.717000",
      "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>,<br>\nthe normal daily routine is different in every hospital, there are houses where the junior radiologists send each and every report to second reading and places where only the \"questionable\" ones are sent to second reader.</p>\n<p>I think that the \"normal\" tagged images come from a screening ( young ppl must have a chest x-ray before starting work to exclude pulmonary TB ) and these are with high % without finding.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1211861,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-20T16:09:58.047000",
          "content": "<p>Thanks, that sounds plausible. It was not clearly described in the paper on the dataset. Largely a good job on those, given that no findings seems to then have mostly been confirmed. But, probably raises the risk that a subtle finding was missed originally and that the reviewers for this dataset got some fatigue after clicking though one\"no finding\" after another and missed it, too, I'd guess?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1211891,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2021-02-20T16:29:04.213000",
          "content": "<p>It is also just guessing from my side… </p>\n<p>Sure, it is always a possibility to have subtle finding on them though!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1184022,
      "author_name": "quillio",
      "author_url": "",
      "post_date": "2021-02-03T11:05:01.830000",
      "content": "<p>Interesting!</p>\n<p>Disregarding the &gt;99% no finding radiologists, it seems unlikely that R8, R9, R10 get no findings 80% of the time.</p>\n<p>Unless the xrays were distributed non-randomly.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1184065,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-03T11:24:54.730000",
          "content": "<p>Yes, I think it's pretty clear that the x-rays were assigned to radiologists in a non-random way. I'm guessing that they split by the original assigned diagnosis at the two hospitals where they got it from (or perhaps they ran some existing algorithm over it) and then somehow systematically assigned \"no findings\" x-rays mostly to one set of reviewers and ones with findings to others (or something like that).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1184083,
          "author_name": "quillio",
          "author_url": "",
          "post_date": "2021-02-03T11:33:09.537000",
          "content": "<p>It's odd that from what I can see, all radiologists agree when there's \"no finding.\"</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1181501": "This [other discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/215444) prompted me to look at the performance of the 17 different radiologists. I found something interesting ([link to notebook](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/#1.-Labelling-process-that-created-the-training-and-test-data)): To my surprise several of them had a 100% match in terms of `class_id` values assigned to each image with the other three radiologists looking a the same image.\n\nIt turns out that radiologists that 100% agree with other radiologists only ever assigned `class_id = 14` (=\"no finding\"). I'd guess X-rays that were a-priori identified (how?) to be without a finding and were sent to a subset of radiologists, especially `R1` to `R7` (and sometimes others, but , `R8`, `R9` and `R10` mostly did not get those images).\n\nNone of this is mentioned in [the paper](https://arxiv.org/pdf/2012.15029.pdf). So I'm wondering whether there's something I overlooked and whether there's other things I did miss about the training process. \n\nIt also makes me think how much we should trust the \"No finding\" images... If someone has to go through 1500 to 3000 images, one after another and there's just never anything there, do they get tired and sloppy (and their colleagues, too)? I would have thought that mixing things up would be a good idea to make sure the radiologists are kept engaged/interested/\"on their toes\" (or however else you want to describe it).\n\n| rad_id | Images with no finding | Percent with no finding |\n| --- | --- |\n|R1 |\t1995 |100% |\n|R2\t| 3118 | 99.9% |\n|R3\t| 2285 | 100% |\n|R4\t| 1513 | 100% |\n|R5\t| 2783 | 100% |\n|R6\t| 2041 | 100% |\n|R7\t| 1733 | 100% |\n|R8\t| 2436 | 19.97% |\n|R9\t| 1979 | 12.6% |\n|R10\t| 2321 | 17.46% |\n|R11 | 1413 | 84.61% |\n|R12|1580 | 91.38% |\n|R13\t|1505 | 82.51% |\n|R14\t|1300 | 80.05% |\n|R15\t|1508 | 82.72% |\n|R16\t|1565 | 88.77% |\n|R17 | 743 | 91.5% |\n",
    "1211830": "@bjoernholzhauer,\nthe normal daily routine is different in every hospital, there are houses where the junior radiologists send each and every report to second reading and places where only the \"questionable\" ones are sent to second reader.\n\nI think that the \"normal\" tagged images come from a screening ( young ppl must have a chest x-ray before starting work to exclude pulmonary TB ) and these are with high % without finding.",
    "1184022": "Interesting!\n\nDisregarding the >99% no finding radiologists, it seems unlikely that R8, R9, R10 get no findings <20% of the time while R11-R17 get no findings >80% of the time.\n\nUnless the xrays were distributed non-randomly."
  }
}