{
  "id": 217886,
  "title": "Treatment of No findings 'Real' vs 'Superficial' No findings.❓❓❓",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/217886",
  "author_name": "sourabhsc",
  "post_date": "2021-02-08T16:06:31.356000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I think out of all the abnormal annotations in training images some annotations would have real 'No findings' while others would just be \"wrongly annotated\" or \"not annotated at all\"  or just 'Superficial' No findings'.<br>\n``Here is the argument </p>\n<ul>\n<li>So All normal annotations = 36096 </li>\n<li>All Abnormal annotations = total annotations - normal annotations 67914- 36096 = 31818</li>\n<li>With the 11606 abnormal annotated images and 3 radiologists annotating the images we get = 11606*3 = 34818 annotations. </li>\n</ul>\n<p>That means there is a discrepancy of 3000 annotations. </p>\n<blockquote>\n  <p>All normal annotations  + abnormal images annotated by 3 radiologists = All abnormal annotations +3000</p>\n</blockquote>\n<p>I think 3000 should be reliable good annotations confirming the 'No findings'. So with 3 radiologists annotating them, I think there are 1000 images with reliable 'No findings' while others are just no annotated properly. </p>\n<p><strong>- WHICH 1000 images???</strong><br>\nFor now, I think one could just pick 1000 images of the 11606 'Abnormal images' randomly and put it in the sample as 15th Class 'real No finding'.</p>\n<ul>\n<li>One needs to check the performance with this extra class.  </li>\n</ul>\n<p>💥 Is this argument correct? What do others think of it? Thank you for your suggestions.</p>\n<p>Check out the vein diagram to explain this :- <br>\n<a href=\"https://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings\" target=\"_blank\">https://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings</a></p>\n<p>PS- Kaggle is not letting me add the image here for some reason. </p>",
  "messages": [
    {
      "id": 1191707,
      "postDate": "2021-02-08T16:06:31.357Z",
      "content": "<p>I think out of all the abnormal annotations in training images some annotations would have real 'No findings' while others would just be \"wrongly annotated\" or \"not annotated at all\"  or just 'Superficial' No findings'.<br>\n``Here is the argument </p>\n<ul>\n<li>So All normal annotations = 36096 </li>\n<li>All Abnormal annotations = total annotations - normal annotations 67914- 36096 = 31818</li>\n<li>With the 11606 abnormal annotated images and 3 radiologists annotating the images we get = 11606*3 = 34818 annotations. </li>\n</ul>\n<p>That means there is a discrepancy of 3000 annotations. </p>\n<blockquote>\n  <p>All normal annotations  + abnormal images annotated by 3 radiologists = All abnormal annotations +3000</p>\n</blockquote>\n<p>I think 3000 should be reliable good annotations confirming the 'No findings'. So with 3 radiologists annotating them, I think there are 1000 images with reliable 'No findings' while others are just no annotated properly. </p>\n<p><strong>- WHICH 1000 images???</strong><br>\nFor now, I think one could just pick 1000 images of the 11606 'Abnormal images' randomly and put it in the sample as 15th Class 'real No finding'.</p>\n<ul>\n<li>One needs to check the performance with this extra class.  </li>\n</ul>\n<p>💥 Is this argument correct? What do others think of it? Thank you for your suggestions.</p>\n<p>Check out the vein diagram to explain this :- <br>\n<a href=\"https://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings\" target=\"_blank\">https://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings</a></p>\n<p>PS- Kaggle is not letting me add the image here for some reason. </p>",
      "rawMarkdown": "I think out of all the abnormal annotations in training images some annotations would have real 'No findings' while others would just be \"wrongly annotated\" or \"not annotated at all\"  or just 'Superficial' No findings'.\n``Here is the argument \n- So All normal annotations = 36096 \n- All Abnormal annotations = total annotations - normal annotations 67914- 36096 = 31818\n- With the 11606 abnormal annotated images and 3 radiologists annotating the images we get = 11606*3 = 34818 annotations. \n\nThat means there is a discrepancy of 3000 annotations. \n> All normal annotations  + abnormal images annotated by 3 radiologists = All abnormal annotations +3000\n\n  I think 3000 should be reliable good annotations confirming the 'No findings'. So with 3 radiologists annotating them, I think there are 1000 images with reliable 'No findings' while others are just no annotated properly. \n\n**- WHICH 1000 images???**\nFor now, I think one could just pick 1000 images of the 11606 'Abnormal images' randomly and put it in the sample as 15th Class 'real No finding'.\n\n- One needs to check the performance with this extra class.  \n\n💥 Is this argument correct? What do others think of it? Thank you for your suggestions.\n\nCheck out the vein diagram to explain this :- \nhttps://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings\n\nPS- Kaggle is not letting me add the image here for some reason. \n\n",
      "votes": 3
    },
    {
      "id": 1197018,
      "postDate": "2021-02-11T21:16:44.503Z",
      "content": "<p>Thank you for your explanation <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> . I was also a bit confused about that jump. That’s why I put this here in the discussion forum so that I can see the bigger picture. I agree with the 5 possibilities of the no findings. <br>\nI am not sure how we can quantify the radiologists’s detection fatigue. </p>\n<p>Your exhaustive discussion on topic will definitely be helpful for me to understand more. I have used the 2 class model and it gives the improvement of 0.079 on public leaderboard for me. But I feel a further sub division of the No finding class in the 2 class methods is also required. I think that one of the method could be incorporated is by creating features that somehow incorporate the different ways a radiologist can mark an X-ray image as ‘No findings’. I am still thinking about how to proceed. <br>\nLet me know if you have some ideas about it. :) </p>",
      "rawMarkdown": "Thank you for your explanation @bjoernholzhauer . I was also a bit confused about that jump. That’s why I put this here in the discussion forum so that I can see the bigger picture. I agree with the 5 possibilities of the no findings. \nI am not sure how we can quantify the radiologists’s detection fatigue. \n\nYour exhaustive discussion on topic will definitely be helpful for me to understand more. I have used the 2 class model and it gives the improvement of 0.079 on public leaderboard for me. But I feel a further sub division of the No finding class in the 2 class methods is also required. I think that one of the method could be incorporated is by creating features that somehow incorporate the different ways a radiologist can mark an X-ray image as ‘No findings’. I am still thinking about how to proceed. \nLet me know if you have some ideas about it. :) ",
      "votes": 1
    },
    {
      "id": 1192654,
      "postDate": "2021-02-09T08:47:27.930Z",
      "content": "<p>I'm not sure it's not, at all, how you jump from 3000 with discrepancies in annotations to 1000 of these images being ones with good annotations of \"no findings\".</p>\n<p>When the three radiologists disagree on whether there's a finding, there are surely multiple scenarios:</p>\n<ul>\n<li>There really should be no findings, but one or two of the radiologists spotted something that made them wrongly draw a box around a finding.</li>\n<li>There really should be no findings, but some radiologist accidentally used the interface wrong and entered a finding.</li>\n<li>There is something there, but one or two of the radiologists missed it.</li>\n<li>There is something there, but some radiologist accidentally did not use the interface correctly and mislabeled the image as \"non-findings\".</li>\n<li>The X-ray is sufficiently inconclusive that even if you ask the radiologists to look at the area of the bounding box, they are going to be unsure and/or disagree with each other what is in that area.</li>\n</ul>\n<p>I'd assume that with a data set this size, you will have a mixture of all five scenarios. </p>\n<p>We've seen evidence for at least some of these scenarios (and the others just seem very plausible to me): We do know there are a bunch of cases that seem like users <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/\" target=\"_blank\">messing up their labels</a> in the user interface (see also <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859\" target=\"_blank\">this discussion</a>). <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">We've seen</a> that the reviewer agreement depends heavily on the case mix, they tend to agree a lot, if they only get fed no-findings cases, but disagree a lot more on a set of cases with findings (60 to 75% agreement). In addition, I suspect - as I (previously posted)[<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181</a>] - that we also have the case of all three radiologists missing a finding partially due to reviewer fatigue when reviewing lots of \"no findings\" images in a row.</p>\n<p>One could think about whether majority voting is an option, or perhaps one should train a model and then see what the model thinks, too. A good model might be a sensible way to resolve blatant errors, but it will be hard for non-radiologists to judge whether such a model is doing the right thing. Note also, that one option for this competition is a two-step approach, where you first classify as \"non-findings\" vs. \"findings - you could see what a model trained for that will do.</p>",
      "rawMarkdown": "I'm not sure it's not, at all, how you jump from 3000 with discrepancies in annotations to 1000 of these images being ones with good annotations of \"no findings\".\n\nWhen the three radiologists disagree on whether there's a finding, there are surely multiple scenarios:\n* There really should be no findings, but one or two of the radiologists spotted something that made them wrongly draw a box around a finding.\n* There really should be no findings, but some radiologist accidentally used the interface wrong and entered a finding.\n* There is something there, but one or two of the radiologists missed it.\n* There is something there, but some radiologist accidentally did not use the interface correctly and mislabeled the image as \"non-findings\".\n* The X-ray is sufficiently inconclusive that even if you ask the radiologists to look at the area of the bounding box, they are going to be unsure and/or disagree with each other what is in that area.\n\nI'd assume that with a data set this size, you will have a mixture of all five scenarios. \n\nWe've seen evidence for at least some of these scenarios (and the others just seem very plausible to me): We do know there are a bunch of cases that seem like users [messing up their labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/) in the user interface (see also [this discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859)). [We've seen](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data) that the reviewer agreement depends heavily on the case mix, they tend to agree a lot, if they only get fed no-findings cases, but disagree a lot more on a set of cases with findings (60 to 75% agreement). In addition, I suspect - as I (previously posted)[https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181] - that we also have the case of all three radiologists missing a finding partially due to reviewer fatigue when reviewing lots of \"no findings\" images in a row.\n\nOne could think about whether majority voting is an option, or perhaps one should train a model and then see what the model thinks, too. A good model might be a sensible way to resolve blatant errors, but it will be hard for non-radiologists to judge whether such a model is doing the right thing. Note also, that one option for this competition is a two-step approach, where you first classify as \"non-findings\" vs. \"findings - you could see what a model trained for that will do.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1197018,
      "author_name": "sourabhsc",
      "author_url": "",
      "post_date": "2021-02-11T21:16:44.503000",
      "content": "<p>Thank you for your explanation <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> . I was also a bit confused about that jump. That’s why I put this here in the discussion forum so that I can see the bigger picture. I agree with the 5 possibilities of the no findings. <br>\nI am not sure how we can quantify the radiologists’s detection fatigue. </p>\n<p>Your exhaustive discussion on topic will definitely be helpful for me to understand more. I have used the 2 class model and it gives the improvement of 0.079 on public leaderboard for me. But I feel a further sub division of the No finding class in the 2 class methods is also required. I think that one of the method could be incorporated is by creating features that somehow incorporate the different ways a radiologist can mark an X-ray image as ‘No findings’. I am still thinking about how to proceed. <br>\nLet me know if you have some ideas about it. :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1192654,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-09T08:47:27.930000",
      "content": "<p>I'm not sure it's not, at all, how you jump from 3000 with discrepancies in annotations to 1000 of these images being ones with good annotations of \"no findings\".</p>\n<p>When the three radiologists disagree on whether there's a finding, there are surely multiple scenarios:</p>\n<ul>\n<li>There really should be no findings, but one or two of the radiologists spotted something that made them wrongly draw a box around a finding.</li>\n<li>There really should be no findings, but some radiologist accidentally used the interface wrong and entered a finding.</li>\n<li>There is something there, but one or two of the radiologists missed it.</li>\n<li>There is something there, but some radiologist accidentally did not use the interface correctly and mislabeled the image as \"non-findings\".</li>\n<li>The X-ray is sufficiently inconclusive that even if you ask the radiologists to look at the area of the bounding box, they are going to be unsure and/or disagree with each other what is in that area.</li>\n</ul>\n<p>I'd assume that with a data set this size, you will have a mixture of all five scenarios. </p>\n<p>We've seen evidence for at least some of these scenarios (and the others just seem very plausible to me): We do know there are a bunch of cases that seem like users <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/\" target=\"_blank\">messing up their labels</a> in the user interface (see also <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859\" target=\"_blank\">this discussion</a>). <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">We've seen</a> that the reviewer agreement depends heavily on the case mix, they tend to agree a lot, if they only get fed no-findings cases, but disagree a lot more on a set of cases with findings (60 to 75% agreement). In addition, I suspect - as I (previously posted)[<a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181\" target=\"_blank\">https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181</a>] - that we also have the case of all three radiologists missing a finding partially due to reviewer fatigue when reviewing lots of \"no findings\" images in a row.</p>\n<p>One could think about whether majority voting is an option, or perhaps one should train a model and then see what the model thinks, too. A good model might be a sensible way to resolve blatant errors, but it will be hard for non-radiologists to judge whether such a model is doing the right thing. Note also, that one option for this competition is a two-step approach, where you first classify as \"non-findings\" vs. \"findings - you could see what a model trained for that will do.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1191707": "I think out of all the abnormal annotations in training images some annotations would have real 'No findings' while others would just be \"wrongly annotated\" or \"not annotated at all\"  or just 'Superficial' No findings'.\n``Here is the argument \n- So All normal annotations = 36096 \n- All Abnormal annotations = total annotations - normal annotations 67914- 36096 = 31818\n- With the 11606 abnormal annotated images and 3 radiologists annotating the images we get = 11606*3 = 34818 annotations. \n\nThat means there is a discrepancy of 3000 annotations. \n> All normal annotations  + abnormal images annotated by 3 radiologists = All abnormal annotations +3000\n\n  I think 3000 should be reliable good annotations confirming the 'No findings'. So with 3 radiologists annotating them, I think there are 1000 images with reliable 'No findings' while others are just no annotated properly. \n\n**- WHICH 1000 images???**\nFor now, I think one could just pick 1000 images of the 11606 'Abnormal images' randomly and put it in the sample as 15th Class 'real No finding'.\n\n- One needs to check the performance with this extra class.  \n\n💥 Is this argument correct? What do others think of it? Thank you for your suggestions.\n\nCheck out the vein diagram to explain this :- \nhttps://www.kaggle.com/sourabhchauhan/vein-diagram-for-3000-real-no-findings\n\nPS- Kaggle is not letting me add the image here for some reason. \n\n",
    "1197018": "Thank you for your explanation @bjoernholzhauer . I was also a bit confused about that jump. That’s why I put this here in the discussion forum so that I can see the bigger picture. I agree with the 5 possibilities of the no findings. \nI am not sure how we can quantify the radiologists’s detection fatigue. \n\nYour exhaustive discussion on topic will definitely be helpful for me to understand more. I have used the 2 class model and it gives the improvement of 0.079 on public leaderboard for me. But I feel a further sub division of the No finding class in the 2 class methods is also required. I think that one of the method could be incorporated is by creating features that somehow incorporate the different ways a radiologist can mark an X-ray image as ‘No findings’. I am still thinking about how to proceed. \nLet me know if you have some ideas about it. :) ",
    "1192654": "I'm not sure it's not, at all, how you jump from 3000 with discrepancies in annotations to 1000 of these images being ones with good annotations of \"no findings\".\n\nWhen the three radiologists disagree on whether there's a finding, there are surely multiple scenarios:\n* There really should be no findings, but one or two of the radiologists spotted something that made them wrongly draw a box around a finding.\n* There really should be no findings, but some radiologist accidentally used the interface wrong and entered a finding.\n* There is something there, but one or two of the radiologists missed it.\n* There is something there, but some radiologist accidentally did not use the interface correctly and mislabeled the image as \"non-findings\".\n* The X-ray is sufficiently inconclusive that even if you ask the radiologists to look at the area of the bounding box, they are going to be unsure and/or disagree with each other what is in that area.\n\nI'd assume that with a data set this size, you will have a mixture of all five scenarios. \n\nWe've seen evidence for at least some of these scenarios (and the others just seem very plausible to me): We do know there are a bunch of cases that seem like users [messing up their labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/) in the user interface (see also [this discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/212859)). [We've seen](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data) that the reviewer agreement depends heavily on the case mix, they tend to agree a lot, if they only get fed no-findings cases, but disagree a lot more on a set of cases with findings (60 to 75% agreement). In addition, I suspect - as I (previously posted)[https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181] - that we also have the case of all three radiologists missing a finding partially due to reviewer fatigue when reviewing lots of \"no findings\" images in a row.\n\nOne could think about whether majority voting is an option, or perhaps one should train a model and then see what the model thinks, too. A good model might be a sensible way to resolve blatant errors, but it will be hard for non-radiologists to judge whether such a model is doing the right thing. Note also, that one option for this competition is a two-step approach, where you first classify as \"non-findings\" vs. \"findings - you could see what a model trained for that will do."
  }
}