{
  "id": 215444,
  "title": "What is the role of rad_id?",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/215444",
  "author_name": "Jack Ding",
  "post_date": "2021-01-29T20:35:40.244000",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>When I was looking through the data I noticed that even though there is a \"rad_id\" column in the training data, when it comes to the test data only the image id is given. So I'm assuming that whatever the model we create will not take radiologist id into account at all, right? On the other hand, the competition page says that a key part of this competition is working with data labelled by different radiologists. So I'm just wondering, what is the exact role of the rad_id column? </p>",
  "messages": [
    {
      "id": 1176837,
      "postDate": "2021-01-29T20:35:40.243Z",
      "content": "<p>Hi Everyone,</p>\n<p>When I was looking through the data I noticed that even though there is a \"rad_id\" column in the training data, when it comes to the test data only the image id is given. So I'm assuming that whatever the model we create will not take radiologist id into account at all, right? On the other hand, the competition page says that a key part of this competition is working with data labelled by different radiologists. So I'm just wondering, what is the exact role of the rad_id column? </p>",
      "rawMarkdown": "Hi Everyone,\n\nWhen I was looking through the data I noticed that even though there is a \"rad_id\" column in the training data, when it comes to the test data only the image id is given. So I'm assuming that whatever the model we create will not take radiologist id into account at all, right? On the other hand, the competition page says that a key part of this competition is working with data labelled by different radiologists. So I'm just wondering, what is the exact role of the rad_id column? ",
      "votes": 6
    },
    {
      "id": 1180420,
      "postDate": "2021-02-01T08:40:24.103Z",
      "content": "<p>You are of course right, <code>rad_id</code> will not be available on the test data, but the information is useful for how to store the data and also for model training. Note the test data have a <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">different data generating process</a> than the training data. That motivates why we might want to do some of the things I describe in the second point below.</p>\n<p>Firstly, the training data consist of annotations by three different radiologists for each image (which three out of the pool varies across image). Unless you do an additional step to get a consensus like they did for the test set (description is in the <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">paper linked on the competition overview page</a>), you will have three separate opinions per image. It's helpful to keep these separate: Three radiologists saying that there's something in the same region is not the same thing as one radiologist saying that the same finding is present three times in overlapping bounding boxes. For the purpose of telling these things apart, the <code>rad_id</code> simply provides a unique ID variable that lets you do this.</p>\n<p>Secondly, there might be good ideas for what we might do with this information during training. E.g.</p>\n<ul>\n<li>A lot of notebooks (e.g. <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">this one</a>) look at how to fuse the different radiologist opinions to get a single ground truth for each image. That makes a lot of sense, because we know that annotations/bounding boxes/bounding box labels from individual radiologists tend to <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes\" target=\"_blank\">have errors and do in this competition</a> (after all, that's why the test set has the consensus process to deal - to some extent - with that label noise).</li>\n<li>Mostly, methods for fusing bounding boxes look at how much the different bounding boxes overlap etc., but they don't tend to take into account whether the same radiologist was more reliable/tended to overlap more on other images. Using <code>rad_id</code>, you could try to account for that (not an easy problem, if anyone saw a good proposal for this, please, link). Perhaps some radiologist even tend to make particular mistakes (whether those are misinterpretations or software usage issues that keep happening to the same person).</li>\n<li>Maybe you can even have an end-to-end model that in one go learns to do the detection task, but also assesses the different annotations and learns to appropriately combine them (with sensible discounting of opinions of radiologists that are more often wrong). It's not really obvious what would be a good way of doing this, but it might have the advantage that our model might be able to \"adjudicate\" (by learning from the other images, taking into account <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#2.-First-part-of-the-data:-train.csv\" target=\"_blank\">correlation of different labels</a> or <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?\" target=\"_blank\">location of bounding boxes for particular labels</a>), when one radiologist often disagrees with the other two,  whether that's this person being right a lot when others fail or whether the person is wrong a lot.</li>\n<li>For each image, you could sample a different radiologists annotation for an image (does not really need unique <code>rad_id</code> across different images, you could just have \"annotator #1 for image #10\" instead; <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed#Let's-add-parallel-processing-in-our-DataLoader\" target=\"_blank\">here's a parallelized PyTorch  DataLoader</a> that does that) or even for all images (a bit tricky, because not all radiologists assessed all images). I guess you even only use the assessments from a few radiologists in each epoch (=not all images would necessarily be used). You might hope that this way your model is exposed to a distribution of viewpoints and would somehow find a \"good consensus\".</li>\n</ul>\n<p>I even wonder whether the competition hosts want to motivate additional research into how to best use labels from different radiologists that you have for your training data.</p>",
      "rawMarkdown": "You are of course right, `rad_id` will not be available on the test data, but the information is useful for how to store the data and also for model training. Note the test data have a [different data generating process](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data) than the training data. That motivates why we might want to do some of the things I describe in the second point below.\n\nFirstly, the training data consist of annotations by three different radiologists for each image (which three out of the pool varies across image). Unless you do an additional step to get a consensus like they did for the test set (description is in the [paper linked on the competition overview page](https://arxiv.org/pdf/2012.15029.pdf)), you will have three separate opinions per image. It's helpful to keep these separate: Three radiologists saying that there's something in the same region is not the same thing as one radiologist saying that the same finding is present three times in overlapping bounding boxes. For the purpose of telling these things apart, the `rad_id` simply provides a unique ID variable that lets you do this.\n\nSecondly, there might be good ideas for what we might do with this information during training. E.g.\n* A lot of notebooks (e.g. [this one](https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset)) look at how to fuse the different radiologist opinions to get a single ground truth for each image. That makes a lot of sense, because we know that annotations/bounding boxes/bounding box labels from individual radiologists tend to [have errors and do in this competition](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes) (after all, that's why the test set has the consensus process to deal - to some extent - with that label noise).\n* Mostly, methods for fusing bounding boxes look at how much the different bounding boxes overlap etc., but they don't tend to take into account whether the same radiologist was more reliable/tended to overlap more on other images. Using `rad_id`, you could try to account for that (not an easy problem, if anyone saw a good proposal for this, please, link). Perhaps some radiologist even tend to make particular mistakes (whether those are misinterpretations or software usage issues that keep happening to the same person).\n* Maybe you can even have an end-to-end model that in one go learns to do the detection task, but also assesses the different annotations and learns to appropriately combine them (with sensible discounting of opinions of radiologists that are more often wrong). It's not really obvious what would be a good way of doing this, but it might have the advantage that our model might be able to \"adjudicate\" (by learning from the other images, taking into account [correlation of different labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#2.-First-part-of-the-data:-train.csv) or [location of bounding boxes for particular labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?)), when one radiologist often disagrees with the other two,  whether that's this person being right a lot when others fail or whether the person is wrong a lot.\n* For each image, you could sample a different radiologists annotation for an image (does not really need unique `rad_id` across different images, you could just have \"annotator #1 for image #10\" instead; [here's a parallelized PyTorch  DataLoader](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed#Let's-add-parallel-processing-in-our-DataLoader) that does that) or even for all images (a bit tricky, because not all radiologists assessed all images). I guess you even only use the assessments from a few radiologists in each epoch (=not all images would necessarily be used). You might hope that this way your model is exposed to a distribution of viewpoints and would somehow find a \"good consensus\".\n\nI even wonder whether the competition hosts want to motivate additional research into how to best use labels from different radiologists that you have for your training data.",
      "votes": 1,
      "replies": [
        {
          "id": 1181344,
          "postDate": "2021-02-01T19:51:34.720Z",
          "content": "<p>Hi Björn,</p>\n<p>Thank you for the detailed answer! I realized after posting the question that different radiologists might label the same anomaly multiple times with slightly different bounding boxes each time. </p>\n<p>Curiously, I've also noticed different radiologists label the same class, but with disjoint bounding boxes. This seems to suggest that certain radiologists may have blind spots or that certain radiologists are seeing things that aren't there. Were you able to determine which radiologists made the most errors? This seems hard to accomplish without knowing what the correct labels actually are. Maybe we can, as you said, measure which radiologist overlaps the most with the others, but what's weird to me is that we have no way of knowing how they reached a consensus. Since they bring in 5 experts for the test data and only 3 experts for the train data, it could be that the 3 experts all made the same mistake on the labels and a 4th expert corrected all of them. I think the training data could only tell us how much the 3 experts agree with each other right? </p>\n<p>As for fusing bounding boxes, I think there would be different approaches: You could take their intersection vs taking their mean vs taking the smallest box that contains all of them. I'm not sure yet which is the best approach. But I wonder if we should just take the fusion approach which yields the best evaluation score at the end. </p>",
          "rawMarkdown": "Hi Björn,\n\nThank you for the detailed answer! I realized after posting the question that different radiologists might label the same anomaly multiple times with slightly different bounding boxes each time. \n\nCuriously, I've also noticed different radiologists label the same class, but with disjoint bounding boxes. This seems to suggest that certain radiologists may have blind spots or that certain radiologists are seeing things that aren't there. Were you able to determine which radiologists made the most errors? This seems hard to accomplish without knowing what the correct labels actually are. Maybe we can, as you said, measure which radiologist overlaps the most with the others, but what's weird to me is that we have no way of knowing how they reached a consensus. Since they bring in 5 experts for the test data and only 3 experts for the train data, it could be that the 3 experts all made the same mistake on the labels and a 4th expert corrected all of them. I think the training data could only tell us how much the 3 experts agree with each other right? \n\nAs for fusing bounding boxes, I think there would be different approaches: You could take their intersection vs taking their mean vs taking the smallest box that contains all of them. I'm not sure yet which is the best approach. But I wonder if we should just take the fusion approach which yields the best evaluation score at the end. ",
          "votes": 1
        },
        {
          "id": 1181370,
          "postDate": "2021-02-01T20:32:45Z",
          "content": "<p>I've just tried to figure out what's going on with agreement between the radiologists (see end of Section 1 of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">this notebook</a>). However, this is actually a bit hard to interpret. It looks to me like the assignment of cases to radiologists was not at all random and there's some clusters of radiologists that very often co-reviewed the same cases.</p>\n<p>The biggest issue that lets some radiologists have 100% agreement with their colleagues is that apparently some of them only reviewed cases that were selected to have no findings! For those that reviewed 100% \"no findings\" cases and agreed 100% with their colleagues, we'll not be able to say much (but it's also not important to figure anything out about these). For those with a case mix, we should be able to do something. Really intrigued by this now…</p>",
          "rawMarkdown": "I've just tried to figure out what's going on with agreement between the radiologists (see end of Section 1 of [this notebook](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray)). However, this is actually a bit hard to interpret. It looks to me like the assignment of cases to radiologists was not at all random and there's some clusters of radiologists that very often co-reviewed the same cases.\n\nThe biggest issue that lets some radiologists have 100% agreement with their colleagues is that apparently some of them only reviewed cases that were selected to have no findings! For those that reviewed 100% \"no findings\" cases and agreed 100% with their colleagues, we'll not be able to say much (but it's also not important to figure anything out about these). For those with a case mix, we should be able to do something. Really intrigued by this now..."
        }
      ]
    },
    {
      "id": 1193812,
      "postDate": "2021-02-09T21:42:12.053Z",
      "content": "<p>Speaking purely from a conceptual standpoint (as I have not looked at the competition as of yet), you want to create a model that will take as little bias into account as possible (for example, you want it to learn the material well enough, but not too well that it does not perform well on other generalized data; hence the classical need for splitting up training and testing data). As they are providing images from a variety of sources, you will certainly want to take parts from all of those sources! This will help to provide diversity in your input set.  I believe that them providing you with the radiologist id is there purely just to give you, the data scientist, the sanity check that your data is diverse, and coming from a wide set. </p>\n<p>TLDR; The rad_id is just for you to keep track of what radiologist the image came from. It certainly should not be taken into account in the decision-making process of your model. They emphasized the importance of using different sources to create a diverse and well-generalized model.</p>",
      "rawMarkdown": "Speaking purely from a conceptual standpoint (as I have not looked at the competition as of yet), you want to create a model that will take as little bias into account as possible (for example, you want it to learn the material well enough, but not too well that it does not perform well on other generalized data; hence the classical need for splitting up training and testing data). As they are providing images from a variety of sources, you will certainly want to take parts from all of those sources! This will help to provide diversity in your input set.  I believe that them providing you with the radiologist id is there purely just to give you, the data scientist, the sanity check that your data is diverse, and coming from a wide set. \n\nTLDR; The rad_id is just for you to keep track of what radiologist the image came from. It certainly should not be taken into account in the decision-making process of your model. They emphasized the importance of using different sources to create a diverse and well-generalized model."
    },
    {
      "id": 1181343,
      "postDate": "2021-02-01T19:50:39.447Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1180420,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-01T08:40:24.103000",
      "content": "<p>You are of course right, <code>rad_id</code> will not be available on the test data, but the information is useful for how to store the data and also for model training. Note the test data have a <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data\" target=\"_blank\">different data generating process</a> than the training data. That motivates why we might want to do some of the things I describe in the second point below.</p>\n<p>Firstly, the training data consist of annotations by three different radiologists for each image (which three out of the pool varies across image). Unless you do an additional step to get a consensus like they did for the test set (description is in the <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">paper linked on the competition overview page</a>), you will have three separate opinions per image. It's helpful to keep these separate: Three radiologists saying that there's something in the same region is not the same thing as one radiologist saying that the same finding is present three times in overlapping bounding boxes. For the purpose of telling these things apart, the <code>rad_id</code> simply provides a unique ID variable that lets you do this.</p>\n<p>Secondly, there might be good ideas for what we might do with this information during training. E.g.</p>\n<ul>\n<li>A lot of notebooks (e.g. <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">this one</a>) look at how to fuse the different radiologist opinions to get a single ground truth for each image. That makes a lot of sense, because we know that annotations/bounding boxes/bounding box labels from individual radiologists tend to <a href=\"https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes\" target=\"_blank\">have errors and do in this competition</a> (after all, that's why the test set has the consensus process to deal - to some extent - with that label noise).</li>\n<li>Mostly, methods for fusing bounding boxes look at how much the different bounding boxes overlap etc., but they don't tend to take into account whether the same radiologist was more reliable/tended to overlap more on other images. Using <code>rad_id</code>, you could try to account for that (not an easy problem, if anyone saw a good proposal for this, please, link). Perhaps some radiologist even tend to make particular mistakes (whether those are misinterpretations or software usage issues that keep happening to the same person).</li>\n<li>Maybe you can even have an end-to-end model that in one go learns to do the detection task, but also assesses the different annotations and learns to appropriately combine them (with sensible discounting of opinions of radiologists that are more often wrong). It's not really obvious what would be a good way of doing this, but it might have the advantage that our model might be able to \"adjudicate\" (by learning from the other images, taking into account <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#2.-First-part-of-the-data:-train.csv\" target=\"_blank\">correlation of different labels</a> or <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?\" target=\"_blank\">location of bounding boxes for particular labels</a>), when one radiologist often disagrees with the other two,  whether that's this person being right a lot when others fail or whether the person is wrong a lot.</li>\n<li>For each image, you could sample a different radiologists annotation for an image (does not really need unique <code>rad_id</code> across different images, you could just have \"annotator #1 for image #10\" instead; <a href=\"https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed#Let's-add-parallel-processing-in-our-DataLoader\" target=\"_blank\">here's a parallelized PyTorch  DataLoader</a> that does that) or even for all images (a bit tricky, because not all radiologists assessed all images). I guess you even only use the assessments from a few radiologists in each epoch (=not all images would necessarily be used). You might hope that this way your model is exposed to a distribution of viewpoints and would somehow find a \"good consensus\".</li>\n</ul>\n<p>I even wonder whether the competition hosts want to motivate additional research into how to best use labels from different radiologists that you have for your training data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1181344,
          "author_name": "Jack Ding",
          "author_url": "",
          "post_date": "2021-02-01T19:51:34.720000",
          "content": "<p>Hi Björn,</p>\n<p>Thank you for the detailed answer! I realized after posting the question that different radiologists might label the same anomaly multiple times with slightly different bounding boxes each time. </p>\n<p>Curiously, I've also noticed different radiologists label the same class, but with disjoint bounding boxes. This seems to suggest that certain radiologists may have blind spots or that certain radiologists are seeing things that aren't there. Were you able to determine which radiologists made the most errors? This seems hard to accomplish without knowing what the correct labels actually are. Maybe we can, as you said, measure which radiologist overlaps the most with the others, but what's weird to me is that we have no way of knowing how they reached a consensus. Since they bring in 5 experts for the test data and only 3 experts for the train data, it could be that the 3 experts all made the same mistake on the labels and a 4th expert corrected all of them. I think the training data could only tell us how much the 3 experts agree with each other right? </p>\n<p>As for fusing bounding boxes, I think there would be different approaches: You could take their intersection vs taking their mean vs taking the smallest box that contains all of them. I'm not sure yet which is the best approach. But I wonder if we should just take the fusion approach which yields the best evaluation score at the end. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1181370,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-02-01T20:32:45",
          "content": "<p>I've just tried to figure out what's going on with agreement between the radiologists (see end of Section 1 of <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray\" target=\"_blank\">this notebook</a>). However, this is actually a bit hard to interpret. It looks to me like the assignment of cases to radiologists was not at all random and there's some clusters of radiologists that very often co-reviewed the same cases.</p>\n<p>The biggest issue that lets some radiologists have 100% agreement with their colleagues is that apparently some of them only reviewed cases that were selected to have no findings! For those that reviewed 100% \"no findings\" cases and agreed 100% with their colleagues, we'll not be able to say much (but it's also not important to figure anything out about these). For those with a case mix, we should be able to do something. Really intrigued by this now…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1193812,
      "author_name": "jd",
      "author_url": "",
      "post_date": "2021-02-09T21:42:12.053000",
      "content": "<p>Speaking purely from a conceptual standpoint (as I have not looked at the competition as of yet), you want to create a model that will take as little bias into account as possible (for example, you want it to learn the material well enough, but not too well that it does not perform well on other generalized data; hence the classical need for splitting up training and testing data). As they are providing images from a variety of sources, you will certainly want to take parts from all of those sources! This will help to provide diversity in your input set.  I believe that them providing you with the radiologist id is there purely just to give you, the data scientist, the sanity check that your data is diverse, and coming from a wide set. </p>\n<p>TLDR; The rad_id is just for you to keep track of what radiologist the image came from. It certainly should not be taken into account in the decision-making process of your model. They emphasized the importance of using different sources to create a diverse and well-generalized model.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1181343,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-01T19:50:39.447000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1176837": "Hi Everyone,\n\nWhen I was looking through the data I noticed that even though there is a \"rad_id\" column in the training data, when it comes to the test data only the image id is given. So I'm assuming that whatever the model we create will not take radiologist id into account at all, right? On the other hand, the competition page says that a key part of this competition is working with data labelled by different radiologists. So I'm just wondering, what is the exact role of the rad_id column? ",
    "1180420": "You are of course right, `rad_id` will not be available on the test data, but the information is useful for how to store the data and also for model training. Note the test data have a [different data generating process](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#1.-Labelling-process-that-created-the-training-and-test-data) than the training data. That motivates why we might want to do some of the things I describe in the second point below.\n\nFirstly, the training data consist of annotations by three different radiologists for each image (which three out of the pool varies across image). Unless you do an additional step to get a consensus like they did for the test set (description is in the [paper linked on the competition overview page](https://arxiv.org/pdf/2012.15029.pdf)), you will have three separate opinions per image. It's helpful to keep these separate: Three radiologists saying that there's something in the same region is not the same thing as one radiologist saying that the same finding is present three times in overlapping bounding boxes. For the purpose of telling these things apart, the `rad_id` simply provides a unique ID variable that lets you do this.\n\nSecondly, there might be good ideas for what we might do with this information during training. E.g.\n* A lot of notebooks (e.g. [this one](https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset)) look at how to fuse the different radiologist opinions to get a single ground truth for each image. That makes a lot of sense, because we know that annotations/bounding boxes/bounding box labels from individual radiologists tend to [have errors and do in this competition](https://www.kaggle.com/bjoernholzhauer/finding-data-issues-and-mislabeled-bounding-boxes) (after all, that's why the test set has the consensus process to deal - to some extent - with that label noise).\n* Mostly, methods for fusing bounding boxes look at how much the different bounding boxes overlap etc., but they don't tend to take into account whether the same radiologist was more reliable/tended to overlap more on other images. Using `rad_id`, you could try to account for that (not an easy problem, if anyone saw a good proposal for this, please, link). Perhaps some radiologist even tend to make particular mistakes (whether those are misinterpretations or software usage issues that keep happening to the same person).\n* Maybe you can even have an end-to-end model that in one go learns to do the detection task, but also assesses the different annotations and learns to appropriately combine them (with sensible discounting of opinions of radiologists that are more often wrong). It's not really obvious what would be a good way of doing this, but it might have the advantage that our model might be able to \"adjudicate\" (by learning from the other images, taking into account [correlation of different labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#2.-First-part-of-the-data:-train.csv) or [location of bounding boxes for particular labels](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray#4.-Where-do-the-different-findings-tend-to-be?)), when one radiologist often disagrees with the other two,  whether that's this person being right a lot when others fail or whether the person is wrong a lot.\n* For each image, you could sample a different radiologists annotation for an image (does not really need unique `rad_id` across different images, you could just have \"annotator #1 for image #10\" instead; [here's a parallelized PyTorch  DataLoader](https://www.kaggle.com/bjoernholzhauer/vinbigdata-chest-x-ray-comparing-dataloader-speed#Let's-add-parallel-processing-in-our-DataLoader) that does that) or even for all images (a bit tricky, because not all radiologists assessed all images). I guess you even only use the assessments from a few radiologists in each epoch (=not all images would necessarily be used). You might hope that this way your model is exposed to a distribution of viewpoints and would somehow find a \"good consensus\".\n\nI even wonder whether the competition hosts want to motivate additional research into how to best use labels from different radiologists that you have for your training data.",
    "1193812": "Speaking purely from a conceptual standpoint (as I have not looked at the competition as of yet), you want to create a model that will take as little bias into account as possible (for example, you want it to learn the material well enough, but not too well that it does not perform well on other generalized data; hence the classical need for splitting up training and testing data). As they are providing images from a variety of sources, you will certainly want to take parts from all of those sources! This will help to provide diversity in your input set.  I believe that them providing you with the radiologist id is there purely just to give you, the data scientist, the sanity check that your data is diverse, and coming from a wide set. \n\nTLDR; The rad_id is just for you to keep track of what radiologist the image came from. It certainly should not be taken into account in the decision-making process of your model. They emphasized the importance of using different sources to create a diverse and well-generalized model.",
    "1181343": ""
  }
}