{
  "id": 369262,
  "title": "A Brief Intro to Mammography",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369262",
  "author_name": "Ian Pan",
  "post_date": "2022-11-29T14:04:16.246000",
  "votes": 253,
  "comment_count": 35,
  "views": 0,
  "content": "<h2>What is mammography?</h2>\n<p>Mammography is the imaging modality of choice for breast cancer screening in women. It also plays an important role in the evaluation of other breast diseases in all patients. You can think of mammography as an X-ray for the breast. </p>\n<p>Different associations across the world have different guidelines for when to start and how often a patient should undergo breast cancer screening. The American College of Radiology and American College of Obstetrics and Gynecology recommend beginning <strong>annual</strong> screening at age 40. The United States Preventative Services Tasks Force recommends <strong>biannual</strong> screening at age 50. </p>\n<p>Depending on individual risk factors (BRCA gene mutations, strong family history, personal history of mantle radiation, etc.), screening may be recommended to begin earlier. Approximately 1 in 8 women will develop breast cancer in their lifetime. Screening allows for early detection and treatment. </p>\n<h2>What is a screening mammogram?</h2>\n<p>A screening mammogram is the most common type of mammogram performed. During a screening exam, the patient goes to an imaging center for a scheduled appointment, and the mammogram technologist will perform two standard views: craniocaudal (CC) and mediolateral oblique (MLO). </p>\n<p>During mammography, the breast is put into compression, which can be quite uncomfortable for the patient! For the CC view, the breast is compressed from \"head to toe\" while for the MLO view, the breast is compressed \"side to side\" at an angle. This allows radiologists to localize a finding to a particular area in the breast and to determine what type of abnormality is present. The patient then leaves the imaging center after the appointment and awaits the results, which are required to be delivered within 30 days. </p>\n<h2>What is BI-RADS?</h2>\n<p>The Breast Imaging Reporting and Data System \"provides standardized breast imaging terminology, report organization, assessment structure and a classification system for mammography, ultrasound and MRI of the breast\"(<a href=\"https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads)\" target=\"_blank\">https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads)</a>. It is a standard lexicon used by radiologists that allows for improved peer review and quality assurance mechanisms with the goal of improving patient care. </p>\n<p>BI-RADS is not only used for mammography but also for breast ultrasound and breast MRI. It essentially defines the vocabulary for breast imaging, a summary of which can be found here: <a href=\"https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf\" target=\"_blank\">https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf</a></p>\n<p>For the purposes of this challenge, the most important element of BI-RADS to understand is the <strong>BI-RADS score</strong>.</p>\n<p>There are 7 main BI-RADS scores, or categories:<br>\n0 - Need additional imaging evaluation<br>\n1 - Negative<br>\n2 - Benign<br>\n3 - Probably Benign<br>\n4 - Suspicious<br>\n5 - Highly Suggestive of Malignancy<br>\n6 - Known Biopsy-Proven Malignancy</p>\n<p><strong>Screening mammograms can only be assigned a BI-RADS score of 0, 1, or 2.</strong> This is why the dataset only contains 3 of the BI-RADS categories. The reason why will be clear as we talk about the breast imaging workflow.</p>\n<h2>What is the breast imaging workflow?</h2>\n<p>We talked briefly about screening mammography earlier in this post. If a patient's screening mammogram is given BI-RADS 1 or 2, then they will be due for another screening at whatever interval they and their physician decide. </p>\n<p>If a screening mammogram is given BI-RADS 0, this means the interpreting radiologist saw an abnormality (mass, calcification, asymmetry) that warranted further evaluation. <strong>You cannot diagnose cancer on a screening mammogram.</strong> The patient is then \"called back\" for additional imaging. </p>\n<p>At the next appointment, the patient will undergo a <strong>diagnostic</strong> mammogram. The standard MLO and CC views are once obtained, but typically these exams are performed <strong>online</strong>, which means the patient does not leave until all necessary imaging is obtained. After the standard views are obtained, the radiologist will look at the images and determine if additional, more specific, views are needed. Examples include magnification views to examine calcifications or \"spot compression\" views to see if a suspicious area of tissue goes away if more compression is applied. </p>\n<p>Oftentimes, a breast ultrasound is performed as well for additional evaluation. When the radiologist is satisfied with all the imaging, they will assign a final BI-RADS score, which can vary from 1 to 6. </p>\n<p>BI-RADS 3 indicates a probably benign finding that will be followed for a period of time (usually 6, 12, and 24 months). This means the patient returns for a <strong>diagnostic</strong> appointment at each interval, and the interval is more frequent than regular screening. </p>\n<p>BI-RADS 4 indicates a suspicious finding and is further subdivided into 4A, 4B, and 4C, depending on the level of suspicion that the finding may represent cancer. <strong>BI-RADS 4 means you are recommending a biopsy for tissue diagnosis.</strong> </p>\n<p>BI-RADS 5 indicates a highly suspicious finding- a greater than 95% chance that you believe the finding is cancer. BI-RADS 5 also means you are recommending a biopsy/tissue sampling. In fact, if you are assigning BI-RADS 5, even if the biopsy is negative, you would suspect that something went wrong during the biopsy and recommend repeat biopsy or even surgical excision because you believe the finding to be that suspicious.</p>\n<p>BI-RADS 6 is given when the patient has a known malignancy. These are usually pre-treatment planning cases or cases where the patient may be undergoing chemotherapy for a cancer that cannot be surgically treated. </p>\n<h2>Our task</h2>\n<p>For this task, we are focused on <strong>screening mammograms</strong>, which again means that only scores of <strong>BI-RADS 0, 1, or 2</strong> are possible. One can essentially think of 0 as \"abnormal\" and 1 and 2 as \"normal.\" There can be subjectivity in assigning BI-RADS 1 or 2. For example, if there are stable findings that are almost certainly benign breast cysts, the mammogram may be assigned 1 or 2 depending on the radiologist. </p>\n<p>However, our task in the challenge is to predict <strong>cancer or no cancer</strong>, a binary value. This may be obvious, but not all BI-RADS 0 cases have cancer. The final cancer label will depend on the outcome of the diagnostic mammogram and the biopsy results, if obtained. </p>\n<p>In this post, I wanted to give competitors a sense of what mammography is and how the breast imaging workflow is designed so they can contextualize the challenge task within the broader scope of breast cancer diagnosis and treatment. There is a lot more to mammography, but I hope people find this introduction helpful!</p>",
  "messages": [
    {
      "id": 2048515,
      "postDate": "2022-11-29T14:04:16.247Z",
      "content": "<h2>What is mammography?</h2>\n<p>Mammography is the imaging modality of choice for breast cancer screening in women. It also plays an important role in the evaluation of other breast diseases in all patients. You can think of mammography as an X-ray for the breast. </p>\n<p>Different associations across the world have different guidelines for when to start and how often a patient should undergo breast cancer screening. The American College of Radiology and American College of Obstetrics and Gynecology recommend beginning <strong>annual</strong> screening at age 40. The United States Preventative Services Tasks Force recommends <strong>biannual</strong> screening at age 50. </p>\n<p>Depending on individual risk factors (BRCA gene mutations, strong family history, personal history of mantle radiation, etc.), screening may be recommended to begin earlier. Approximately 1 in 8 women will develop breast cancer in their lifetime. Screening allows for early detection and treatment. </p>\n<h2>What is a screening mammogram?</h2>\n<p>A screening mammogram is the most common type of mammogram performed. During a screening exam, the patient goes to an imaging center for a scheduled appointment, and the mammogram technologist will perform two standard views: craniocaudal (CC) and mediolateral oblique (MLO). </p>\n<p>During mammography, the breast is put into compression, which can be quite uncomfortable for the patient! For the CC view, the breast is compressed from \"head to toe\" while for the MLO view, the breast is compressed \"side to side\" at an angle. This allows radiologists to localize a finding to a particular area in the breast and to determine what type of abnormality is present. The patient then leaves the imaging center after the appointment and awaits the results, which are required to be delivered within 30 days. </p>\n<h2>What is BI-RADS?</h2>\n<p>The Breast Imaging Reporting and Data System \"provides standardized breast imaging terminology, report organization, assessment structure and a classification system for mammography, ultrasound and MRI of the breast\"(<a href=\"https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads)\" target=\"_blank\">https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads)</a>. It is a standard lexicon used by radiologists that allows for improved peer review and quality assurance mechanisms with the goal of improving patient care. </p>\n<p>BI-RADS is not only used for mammography but also for breast ultrasound and breast MRI. It essentially defines the vocabulary for breast imaging, a summary of which can be found here: <a href=\"https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf\" target=\"_blank\">https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf</a></p>\n<p>For the purposes of this challenge, the most important element of BI-RADS to understand is the <strong>BI-RADS score</strong>.</p>\n<p>There are 7 main BI-RADS scores, or categories:<br>\n0 - Need additional imaging evaluation<br>\n1 - Negative<br>\n2 - Benign<br>\n3 - Probably Benign<br>\n4 - Suspicious<br>\n5 - Highly Suggestive of Malignancy<br>\n6 - Known Biopsy-Proven Malignancy</p>\n<p><strong>Screening mammograms can only be assigned a BI-RADS score of 0, 1, or 2.</strong> This is why the dataset only contains 3 of the BI-RADS categories. The reason why will be clear as we talk about the breast imaging workflow.</p>\n<h2>What is the breast imaging workflow?</h2>\n<p>We talked briefly about screening mammography earlier in this post. If a patient's screening mammogram is given BI-RADS 1 or 2, then they will be due for another screening at whatever interval they and their physician decide. </p>\n<p>If a screening mammogram is given BI-RADS 0, this means the interpreting radiologist saw an abnormality (mass, calcification, asymmetry) that warranted further evaluation. <strong>You cannot diagnose cancer on a screening mammogram.</strong> The patient is then \"called back\" for additional imaging. </p>\n<p>At the next appointment, the patient will undergo a <strong>diagnostic</strong> mammogram. The standard MLO and CC views are once obtained, but typically these exams are performed <strong>online</strong>, which means the patient does not leave until all necessary imaging is obtained. After the standard views are obtained, the radiologist will look at the images and determine if additional, more specific, views are needed. Examples include magnification views to examine calcifications or \"spot compression\" views to see if a suspicious area of tissue goes away if more compression is applied. </p>\n<p>Oftentimes, a breast ultrasound is performed as well for additional evaluation. When the radiologist is satisfied with all the imaging, they will assign a final BI-RADS score, which can vary from 1 to 6. </p>\n<p>BI-RADS 3 indicates a probably benign finding that will be followed for a period of time (usually 6, 12, and 24 months). This means the patient returns for a <strong>diagnostic</strong> appointment at each interval, and the interval is more frequent than regular screening. </p>\n<p>BI-RADS 4 indicates a suspicious finding and is further subdivided into 4A, 4B, and 4C, depending on the level of suspicion that the finding may represent cancer. <strong>BI-RADS 4 means you are recommending a biopsy for tissue diagnosis.</strong> </p>\n<p>BI-RADS 5 indicates a highly suspicious finding- a greater than 95% chance that you believe the finding is cancer. BI-RADS 5 also means you are recommending a biopsy/tissue sampling. In fact, if you are assigning BI-RADS 5, even if the biopsy is negative, you would suspect that something went wrong during the biopsy and recommend repeat biopsy or even surgical excision because you believe the finding to be that suspicious.</p>\n<p>BI-RADS 6 is given when the patient has a known malignancy. These are usually pre-treatment planning cases or cases where the patient may be undergoing chemotherapy for a cancer that cannot be surgically treated. </p>\n<h2>Our task</h2>\n<p>For this task, we are focused on <strong>screening mammograms</strong>, which again means that only scores of <strong>BI-RADS 0, 1, or 2</strong> are possible. One can essentially think of 0 as \"abnormal\" and 1 and 2 as \"normal.\" There can be subjectivity in assigning BI-RADS 1 or 2. For example, if there are stable findings that are almost certainly benign breast cysts, the mammogram may be assigned 1 or 2 depending on the radiologist. </p>\n<p>However, our task in the challenge is to predict <strong>cancer or no cancer</strong>, a binary value. This may be obvious, but not all BI-RADS 0 cases have cancer. The final cancer label will depend on the outcome of the diagnostic mammogram and the biopsy results, if obtained. </p>\n<p>In this post, I wanted to give competitors a sense of what mammography is and how the breast imaging workflow is designed so they can contextualize the challenge task within the broader scope of breast cancer diagnosis and treatment. There is a lot more to mammography, but I hope people find this introduction helpful!</p>",
      "rawMarkdown": "## What is mammography?\n\nMammography is the imaging modality of choice for breast cancer screening in women. It also plays an important role in the evaluation of other breast diseases in all patients. You can think of mammography as an X-ray for the breast. \n\nDifferent associations across the world have different guidelines for when to start and how often a patient should undergo breast cancer screening. The American College of Radiology and American College of Obstetrics and Gynecology recommend beginning **annual** screening at age 40. The United States Preventative Services Tasks Force recommends **biannual** screening at age 50. \n\nDepending on individual risk factors (BRCA gene mutations, strong family history, personal history of mantle radiation, etc.), screening may be recommended to begin earlier. Approximately 1 in 8 women will develop breast cancer in their lifetime. Screening allows for early detection and treatment. \n\n## What is a screening mammogram?\n\nA screening mammogram is the most common type of mammogram performed. During a screening exam, the patient goes to an imaging center for a scheduled appointment, and the mammogram technologist will perform two standard views: craniocaudal (CC) and mediolateral oblique (MLO). \n\nDuring mammography, the breast is put into compression, which can be quite uncomfortable for the patient! For the CC view, the breast is compressed from \"head to toe\" while for the MLO view, the breast is compressed \"side to side\" at an angle. This allows radiologists to localize a finding to a particular area in the breast and to determine what type of abnormality is present. The patient then leaves the imaging center after the appointment and awaits the results, which are required to be delivered within 30 days. \n\n## What is BI-RADS?\n\nThe Breast Imaging Reporting and Data System \"provides standardized breast imaging terminology, report organization, assessment structure and a classification system for mammography, ultrasound and MRI of the breast\"(https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads). It is a standard lexicon used by radiologists that allows for improved peer review and quality assurance mechanisms with the goal of improving patient care. \n\nBI-RADS is not only used for mammography but also for breast ultrasound and breast MRI. It essentially defines the vocabulary for breast imaging, a summary of which can be found here: https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf\n\nFor the purposes of this challenge, the most important element of BI-RADS to understand is the **BI-RADS score**.\n\nThere are 7 main BI-RADS scores, or categories:\n0 - Need additional imaging evaluation\n1 - Negative\n2 - Benign\n3 - Probably Benign\n4 - Suspicious\n5 - Highly Suggestive of Malignancy\n6 - Known Biopsy-Proven Malignancy\n\n**Screening mammograms can only be assigned a BI-RADS score of 0, 1, or 2.** This is why the dataset only contains 3 of the BI-RADS categories. The reason why will be clear as we talk about the breast imaging workflow.\n\n## What is the breast imaging workflow?\n\nWe talked briefly about screening mammography earlier in this post. If a patient's screening mammogram is given BI-RADS 1 or 2, then they will be due for another screening at whatever interval they and their physician decide. \n\nIf a screening mammogram is given BI-RADS 0, this means the interpreting radiologist saw an abnormality (mass, calcification, asymmetry) that warranted further evaluation. **You cannot diagnose cancer on a screening mammogram.** The patient is then \"called back\" for additional imaging. \n\nAt the next appointment, the patient will undergo a **diagnostic** mammogram. The standard MLO and CC views are once obtained, but typically these exams are performed **online**, which means the patient does not leave until all necessary imaging is obtained. After the standard views are obtained, the radiologist will look at the images and determine if additional, more specific, views are needed. Examples include magnification views to examine calcifications or \"spot compression\" views to see if a suspicious area of tissue goes away if more compression is applied. \n\nOftentimes, a breast ultrasound is performed as well for additional evaluation. When the radiologist is satisfied with all the imaging, they will assign a final BI-RADS score, which can vary from 1 to 6. \n\nBI-RADS 3 indicates a probably benign finding that will be followed for a period of time (usually 6, 12, and 24 months). This means the patient returns for a **diagnostic** appointment at each interval, and the interval is more frequent than regular screening. \n\nBI-RADS 4 indicates a suspicious finding and is further subdivided into 4A, 4B, and 4C, depending on the level of suspicion that the finding may represent cancer. **BI-RADS 4 means you are recommending a biopsy for tissue diagnosis.** \n\nBI-RADS 5 indicates a highly suspicious finding- a greater than 95% chance that you believe the finding is cancer. BI-RADS 5 also means you are recommending a biopsy/tissue sampling. In fact, if you are assigning BI-RADS 5, even if the biopsy is negative, you would suspect that something went wrong during the biopsy and recommend repeat biopsy or even surgical excision because you believe the finding to be that suspicious.\n\nBI-RADS 6 is given when the patient has a known malignancy. These are usually pre-treatment planning cases or cases where the patient may be undergoing chemotherapy for a cancer that cannot be surgically treated. \n\n## Our task\n\nFor this task, we are focused on **screening mammograms**, which again means that only scores of **BI-RADS 0, 1, or 2** are possible. One can essentially think of 0 as \"abnormal\" and 1 and 2 as \"normal.\" There can be subjectivity in assigning BI-RADS 1 or 2. For example, if there are stable findings that are almost certainly benign breast cysts, the mammogram may be assigned 1 or 2 depending on the radiologist. \n\nHowever, our task in the challenge is to predict **cancer or no cancer**, a binary value. This may be obvious, but not all BI-RADS 0 cases have cancer. The final cancer label will depend on the outcome of the diagnostic mammogram and the biopsy results, if obtained. \n\nIn this post, I wanted to give competitors a sense of what mammography is and how the breast imaging workflow is designed so they can contextualize the challenge task within the broader scope of breast cancer diagnosis and treatment. There is a lot more to mammography, but I hope people find this introduction helpful!",
      "votes": 253
    },
    {
      "id": 2052881,
      "postDate": "2022-12-02T15:49:05.210Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thank you for this useful information.</p>\n<p>So if someone gets a BI-RADS of 1 or 2, there will not be any additional actions in the near future. This means that no biopsy will be performed and so there is no chance that the images will be labeled as cancer.</p>\n<p>Do you know, in general, what is the False Negative rate for patients with BI-RADS 1 and 2 for an average physician? Is it around 0.01%, 0.1%, 1%? </p>",
      "rawMarkdown": "@vaillant thank you for this useful information.\n\nSo if someone gets a BI-RADS of 1 or 2, there will not be any additional actions in the near future. This means that no biopsy will be performed and so there is no chance that the images will be labeled as cancer.\n\nDo you know, in general, what is the False Negative rate for patients with BI-RADS 1 and 2 for an average physician? Is it around 0.01%, 0.1%, 1%? ",
      "votes": 3,
      "replies": [
        {
          "id": 2054869,
          "postDate": "2022-12-04T14:04:28.570Z",
          "content": "<p>You are correct. If BI-RADS 1 or 2 is assigned, the only additional action is the the next routine screening. No biopsy is performed, and the image cannot be labeled as having cancer.</p>\n<p>The only caveat is, sometimes cancer will present on a later screening mammogram, and retrospectively one might seen it on a prior screening mammogram initially called 1 or 2. This is rare but does happen. </p>\n<p>There was a study performed in 2012 that found a 20% missed cancer rate (<a href=\"https://pubmed.ncbi.nlm.nih.gov/22700555/)\" target=\"_blank\">https://pubmed.ncbi.nlm.nih.gov/22700555/)</a>, that is, 20% of cancers present at the time of screening were missed on the screening mammogram (thus, screening mammography has a sensitivity about 80%). Certain types of breast cancers (e.g., invasive lobular carcinoma) are harder to detect due to more subtle imaging findings. Fortunately, the most common type of breast cancer (invasive ductal carcinoma) is more easily detectable by imaging. </p>\n<p>That number has likely decreased over the years due to the use of digital breast tomosynthesis (DBT), which can be thought of as a \"CT of the breast.\" This challenge does not use DBT images, just the 2D \"full field\" images. Many breast imaging centers still rely primarily on 2D mammography.</p>",
          "rawMarkdown": "You are correct. If BI-RADS 1 or 2 is assigned, the only additional action is the the next routine screening. No biopsy is performed, and the image cannot be labeled as having cancer.\n\nThe only caveat is, sometimes cancer will present on a later screening mammogram, and retrospectively one might seen it on a prior screening mammogram initially called 1 or 2. This is rare but does happen. \n\nThere was a study performed in 2012 that found a 20% missed cancer rate (https://pubmed.ncbi.nlm.nih.gov/22700555/), that is, 20% of cancers present at the time of screening were missed on the screening mammogram (thus, screening mammography has a sensitivity about 80%). Certain types of breast cancers (e.g., invasive lobular carcinoma) are harder to detect due to more subtle imaging findings. Fortunately, the most common type of breast cancer (invasive ductal carcinoma) is more easily detectable by imaging. \n\nThat number has likely decreased over the years due to the use of digital breast tomosynthesis (DBT), which can be thought of as a \"CT of the breast.\" This challenge does not use DBT images, just the 2D \"full field\" images. Many breast imaging centers still rely primarily on 2D mammography.",
          "votes": 6
        },
        {
          "id": 2054899,
          "postDate": "2022-12-04T14:40:40.013Z",
          "content": "<p>Thanks! Does that mean that there's a potential 20% of images labeled as no cancer that actually contains a cancer in it?</p>\n<p>Or is there a safety procedure where negative images are only labeled after a second exam few years later which is still negative ?</p>\n<p>I know you might not have the answer for this specific dataset but if any of the competition's host could answer that will be great! <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> would you mind passing the question to the hosts please?</p>",
          "rawMarkdown": "Thanks! Does that mean that there's a potential 20% of images labeled as no cancer that actually contains a cancer in it?\n\nOr is there a safety procedure where negative images are only labeled after a second exam few years later which is still negative ?\n\nI know you might not have the answer for this specific dataset but if any of the competition's host could answer that will be great! @inversion would you mind passing the question to the hosts please?",
          "votes": 4
        },
        {
          "id": 2055391,
          "postDate": "2022-12-05T03:54:14.153Z",
          "content": "<p>This may or may not be a very critical question.  On one hand, over diagnosing cancer might be something folks wish to avoid (don't shoot me, I'm not a doctor, so I could be totally way off base there) and so we should find ourselves training <em>not</em> to find cancer.   </p>\n<p>On the other hand, it's a real issue and we want to detect all cancers, no matter how subtle, and the competition hosts have been careful to weed out the false negatives.</p>\n<p>On the third hand =) there are false negatives in the train/test and it's just not a perfect world.  Which goes to back to the first strategy, which is training not to find cancer… </p>\n<p>On the fourth, and frankly most likely hand, this is simply just not a problem to even think about.  First we actually have to get something working.</p>",
          "rawMarkdown": "This may or may not be a very critical question.  On one hand, over diagnosing cancer might be something folks wish to avoid (don't shoot me, I'm not a doctor, so I could be totally way off base there) and so we should find ourselves training *not* to find cancer.   \n\nOn the other hand, it's a real issue and we want to detect all cancers, no matter how subtle, and the competition hosts have been careful to weed out the false negatives.\n\nOn the third hand =) there are false negatives in the train/test and it's just not a perfect world.  Which goes to back to the first strategy, which is training not to find cancer... \n\nOn the fourth, and frankly most likely hand, this is simply just not a problem to even think about.  First we actually have to get something working."
        },
        {
          "id": 2055727,
          "postDate": "2022-12-05T11:08:24.550Z",
          "content": "<p>Whether we are mainly focusing on predicting cancer or no cancer does not change anything in a binary setting, in the end you can play with the threshold of your model to improve either sensitivity or specificity depending on how you intend to use the model.</p>\n<p>In a competition setting, knowing how trustful the labels are might help you go from a somehow working solution to a competitive one.</p>\n<p>In real life setting, which is the important setting, you do not want your model to mimic the current flaws and biases of the average doctor. What you want is the best possible model to detect cancer.</p>\n<p>I am not saying that it is easy or even possible to get a perfect labeling scheme. I am sure the hosts thought about this carefully and created the best possible dataset, nevertheless it's important to know the potential errors that can be hidden in the labeling process, so that maybe someone finds a way to efficiently deal with them. That's why I'd like to know the selection process for negative samples! Or as you say, how did they weed out the false negatives?</p>",
          "rawMarkdown": "Whether we are mainly focusing on predicting cancer or no cancer does not change anything in a binary setting, in the end you can play with the threshold of your model to improve either sensitivity or specificity depending on how you intend to use the model.\n\nIn a competition setting, knowing how trustful the labels are might help you go from a somehow working solution to a competitive one.\n\nIn real life setting, which is the important setting, you do not want your model to mimic the current flaws and biases of the average doctor. What you want is the best possible model to detect cancer.\n\nI am not saying that it is easy or even possible to get a perfect labeling scheme. I am sure the hosts thought about this carefully and created the best possible dataset, nevertheless it's important to know the potential errors that can be hidden in the labeling process, so that maybe someone finds a way to efficiently deal with them. That's why I'd like to know the selection process for negative samples! Or as you say, how did they weed out the false negatives?",
          "votes": 1
        },
        {
          "id": 2056107,
          "postDate": "2022-12-05T18:28:57.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> there's no way to completely eliminate false negatives, but in this case I would expect them to be a very minor issue. Taking the 20% false negative rate as a given, you might expect to see roughly 8,000 patients * 0.2 false negative rate *0.02 cancer rate = 32 false negatives. However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>",
          "rawMarkdown": "@optimo there's no way to completely eliminate false negatives, but in this case I would expect them to be a very minor issue. Taking the 20% false negative rate as a given, you might expect to see roughly 8,000 patients * 0.2 false negative rate *0.02 cancer rate = 32 false negatives. However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.",
          "votes": 7
        },
        {
          "id": 2056609,
          "postDate": "2022-12-06T09:37:27.537Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> ! Very clear, I guess we don't have to worry too much about false negative labels then!</p>",
          "rawMarkdown": "Thank you @sohier ! Very clear, I guess we don't have to worry too much about false negative labels then!"
        },
        {
          "id": 2056790,
          "postDate": "2022-12-06T12:36:38.020Z",
          "content": "<p>Completely agreeing with you! This confusion about false negative cases has puzzled me for a long time. Glad to see this question and answer here.</p>",
          "rawMarkdown": "Completely agreeing with you! This confusion about false negative cases has puzzled me for a long time. Glad to see this question and answer here."
        }
      ]
    },
    {
      "id": 2057112,
      "postDate": "2022-12-06T18:53:46.510Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thanks for very nice write up.</p>\n<p>I am thinking if there is any way to use external data out there (i.e. vindr) which have BI-RADS in 1-5. <br>\nI am thinking of the following logic: <br>\nBI-RADS 4/5 (malignant) := BI-RADS 0 + biopsy<br>\nBI-RADS 2/3 (benign) := BI-RADS 2 or BI-RADS 0 + no-biopsy<br>\nBI-RADS 1 := BI-RADS 1 or BI-RADS is null</p>\n<p>Does this make sense at all?</p>",
      "rawMarkdown": "@vaillant thanks for very nice write up.\n\nI am thinking if there is any way to use external data out there (i.e. vindr) which have BI-RADS in 1-5. \nI am thinking of the following logic: \nBI-RADS 4/5 (malignant) := BI-RADS 0 + biopsy\nBI-RADS 2/3 (benign) := BI-RADS 2 or BI-RADS 0 + no-biopsy\nBI-RADS 1 := BI-RADS 1 or BI-RADS is null\n\nDoes this make sense at all?",
      "votes": 4,
      "replies": [
        {
          "id": 2066075,
          "postDate": "2022-12-15T11:40:44.533Z",
          "content": "<p>I think it could be helpful, but most BI-RADS 4 are benign. There are actually further categories (4A, 4B, 4C) based on level of suspicion. Anything that has 2% or greater probability of cancer is recommended for biopsy.</p>\n<p>If the category 4 subtypes are provided, you could throw 4A into benign and 4C into malignant, ignoring 4B. Not perfect but would be better than just assuming 4 is malignant. </p>\n<p>On another note, I'm not sure if the VinDr-Mammo dataset (<a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a>) is allowed. The data use agreement seems to specify that the data can only be used for research, which goes against the rules of the competition stating that there can be nothing limiting commercial use for the winning solutions. </p>\n<p>Maybe <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can comment.</p>",
          "rawMarkdown": "I think it could be helpful, but most BI-RADS 4 are benign. There are actually further categories (4A, 4B, 4C) based on level of suspicion. Anything that has 2% or greater probability of cancer is recommended for biopsy.\n\nIf the category 4 subtypes are provided, you could throw 4A into benign and 4C into malignant, ignoring 4B. Not perfect but would be better than just assuming 4 is malignant. \n\nOn another note, I'm not sure if the VinDr-Mammo dataset (https://physionet.org/content/vindr-mammo/1.0.0/) is allowed. The data use agreement seems to specify that the data can only be used for research, which goes against the rules of the competition stating that there can be nothing limiting commercial use for the winning solutions. \n\nMaybe @sohier can comment.",
          "votes": 6
        }
      ]
    },
    {
      "id": 2147040,
      "postDate": "2023-02-16T11:23:49.753Z",
      "content": "<p>Thanks for your introduction. Do you think it is reasonable to detect cancer instead abnormalities in Mammography ? </p>",
      "rawMarkdown": "Thanks for your introduction. Do you think it is reasonable to detect cancer instead abnormalities in Mammography ? "
    },
    {
      "id": 2131870,
      "postDate": "2023-02-06T13:00:19.070Z",
      "content": "<p>Thank you for sharing this valuable post!</p>",
      "rawMarkdown": "Thank you for sharing this valuable post!"
    },
    {
      "id": 2084389,
      "postDate": "2023-01-03T13:09:19.047Z",
      "content": "<p>Thanks for your intro, it's great useful!</p>",
      "rawMarkdown": "Thanks for your intro, it's great useful!"
    },
    {
      "id": 2057224,
      "postDate": "2022-12-06T21:32:08.363Z",
      "content": "<p>It is worth reading and know the BTS of the process. This will definitely help us to build the model and understand the data well. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "It is worth reading and know the BTS of the process. This will definitely help us to build the model and understand the data well. @vaillant "
    },
    {
      "id": 2056551,
      "postDate": "2022-12-06T08:21:14.337Z",
      "content": "<p>Thank you for this useful information, I didn't know about any of this… Will you be open to taking me on your team to learn from you, I have some knowledge in computer vision.</p>",
      "rawMarkdown": "Thank you for this useful information, I didn't know about any of this... Will you be open to taking me on your team to learn from you, I have some knowledge in computer vision."
    },
    {
      "id": 2052499,
      "postDate": "2022-12-02T08:56:12.917Z",
      "content": "<p>Thanks for the information, but I have a question about the BI-RADS score. When I look at train.csv, I see that many patients do not have a BIRADS score. Among them are many patients who have cancer. Should I consider this as \"overlooked cancer\" or is there another reason, e.g. privacy?</p>",
      "rawMarkdown": "Thanks for the information, but I have a question about the BI-RADS score. When I look at train.csv, I see that many patients do not have a BIRADS score. Among them are many patients who have cancer. Should I consider this as \"overlooked cancer\" or is there another reason, e.g. privacy?",
      "replies": [
        {
          "id": 2052521,
          "postDate": "2022-12-02T09:24:53.547Z",
          "content": "<p>I think it's not a big due. Maybe some patients have no BIRADS score because they didn't record or share. Through the data we could find out,when the difficult_negetive_case is False and there is a BIRADS score,it must be BIRADS 0 for cancer 1, BIRADS 1 or 2 for cancer 0.If there isn't BIRADS score, We can assume that it follows the above rules.<br>\nIf you have any different opinions, I‘m glad to discuss it with you .</p>",
          "rawMarkdown": "I think it's not a big due. Maybe some patients have no BIRADS score because they didn't record or share. Through the data we could find out,when the difficult_negetive_case is False and there is a BIRADS score,it must be BIRADS 0 for cancer 1, BIRADS 1 or 2 for cancer 0.If there isn't BIRADS score, We can assume that it follows the above rules.\nIf you have any different opinions, I‘m glad to discuss it with you ."
        },
        {
          "id": 2055394,
          "postDate": "2022-12-05T03:57:14.533Z",
          "content": "<p>site2 rarely reported birads.</p>",
          "rawMarkdown": "site2 rarely reported birads."
        }
      ]
    },
    {
      "id": 2052353,
      "postDate": "2022-12-02T05:53:23.923Z",
      "content": "<p>Thanks for sharing valuable information.</p>",
      "rawMarkdown": "Thanks for sharing valuable information."
    },
    {
      "id": 2052325,
      "postDate": "2022-12-02T05:16:32.340Z",
      "content": "<p>Informational. Thanks for explaining briefly</p>",
      "rawMarkdown": "Informational. Thanks for explaining briefly"
    },
    {
      "id": 2050575,
      "postDate": "2022-11-30T21:04:09.860Z",
      "content": "<p>This is incredibly helpful and clear. Thank you!</p>",
      "rawMarkdown": "This is incredibly helpful and clear. Thank you!"
    },
    {
      "id": 2048910,
      "postDate": "2022-11-29T19:16:29.097Z",
      "content": "<p>Very informative!</p>",
      "rawMarkdown": "Very informative!"
    },
    {
      "id": 2048638,
      "postDate": "2022-11-29T15:30:41.607Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for explaining mammography.</p>",
      "rawMarkdown": "Thanks @vaillant for explaining mammography."
    },
    {
      "id": 2049102,
      "postDate": "2022-11-30T00:06:06.457Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2161461,
      "postDate": "2023-02-27T14:14:44.283Z",
      "content": "<p>Thank you for this useful info!</p>",
      "rawMarkdown": "Thank you for this useful info!"
    },
    {
      "id": 2151666,
      "postDate": "2023-02-20T08:58:09.060Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 2097281,
      "postDate": "2023-01-12T14:53:02.390Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    },
    {
      "id": 2085355,
      "postDate": "2023-01-04T05:10:38.930Z",
      "content": "<p>Thank you for useful information!</p>",
      "rawMarkdown": "Thank you for useful information!"
    },
    {
      "id": 2065928,
      "postDate": "2022-12-15T07:56:26.427Z",
      "content": "<p>Thanks for your great introduction. </p>",
      "rawMarkdown": "Thanks for your great introduction. "
    },
    {
      "id": 2054933,
      "postDate": "2022-12-04T15:10:15.503Z",
      "content": "<p>thanks for sharing , very useful </p>",
      "rawMarkdown": "thanks for sharing , very useful "
    },
    {
      "id": 2051546,
      "postDate": "2022-12-01T13:45:25.203Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 2050451,
      "postDate": "2022-11-30T18:37:20.657Z",
      "content": "<p>Thanks for the helpful details</p>",
      "rawMarkdown": "Thanks for the helpful details"
    },
    {
      "id": 2050368,
      "postDate": "2022-11-30T17:35:52.647Z",
      "content": "<p>Thanks for the valuable info mate!</p>",
      "rawMarkdown": "Thanks for the valuable info mate!"
    },
    {
      "id": 2049354,
      "postDate": "2022-11-30T04:31:26.693Z",
      "content": "<p>Thanks for sharing~~~</p>",
      "rawMarkdown": "Thanks for sharing~~~"
    },
    {
      "id": 2049334,
      "postDate": "2022-11-30T04:01:28.990Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2052881,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2022-12-02T15:49:05.210000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thank you for this useful information.</p>\n<p>So if someone gets a BI-RADS of 1 or 2, there will not be any additional actions in the near future. This means that no biopsy will be performed and so there is no chance that the images will be labeled as cancer.</p>\n<p>Do you know, in general, what is the False Negative rate for patients with BI-RADS 1 and 2 for an average physician? Is it around 0.01%, 0.1%, 1%? </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2054869,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-12-04T14:04:28.570000",
          "content": "<p>You are correct. If BI-RADS 1 or 2 is assigned, the only additional action is the the next routine screening. No biopsy is performed, and the image cannot be labeled as having cancer.</p>\n<p>The only caveat is, sometimes cancer will present on a later screening mammogram, and retrospectively one might seen it on a prior screening mammogram initially called 1 or 2. This is rare but does happen. </p>\n<p>There was a study performed in 2012 that found a 20% missed cancer rate (<a href=\"https://pubmed.ncbi.nlm.nih.gov/22700555/)\" target=\"_blank\">https://pubmed.ncbi.nlm.nih.gov/22700555/)</a>, that is, 20% of cancers present at the time of screening were missed on the screening mammogram (thus, screening mammography has a sensitivity about 80%). Certain types of breast cancers (e.g., invasive lobular carcinoma) are harder to detect due to more subtle imaging findings. Fortunately, the most common type of breast cancer (invasive ductal carcinoma) is more easily detectable by imaging. </p>\n<p>That number has likely decreased over the years due to the use of digital breast tomosynthesis (DBT), which can be thought of as a \"CT of the breast.\" This challenge does not use DBT images, just the 2D \"full field\" images. Many breast imaging centers still rely primarily on 2D mammography.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 2054899,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-12-04T14:40:40.013000",
          "content": "<p>Thanks! Does that mean that there's a potential 20% of images labeled as no cancer that actually contains a cancer in it?</p>\n<p>Or is there a safety procedure where negative images are only labeled after a second exam few years later which is still negative ?</p>\n<p>I know you might not have the answer for this specific dataset but if any of the competition's host could answer that will be great! <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> would you mind passing the question to the hosts please?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2055391,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T03:54:14.153000",
          "content": "<p>This may or may not be a very critical question.  On one hand, over diagnosing cancer might be something folks wish to avoid (don't shoot me, I'm not a doctor, so I could be totally way off base there) and so we should find ourselves training <em>not</em> to find cancer.   </p>\n<p>On the other hand, it's a real issue and we want to detect all cancers, no matter how subtle, and the competition hosts have been careful to weed out the false negatives.</p>\n<p>On the third hand =) there are false negatives in the train/test and it's just not a perfect world.  Which goes to back to the first strategy, which is training not to find cancer… </p>\n<p>On the fourth, and frankly most likely hand, this is simply just not a problem to even think about.  First we actually have to get something working.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055727,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-12-05T11:08:24.550000",
          "content": "<p>Whether we are mainly focusing on predicting cancer or no cancer does not change anything in a binary setting, in the end you can play with the threshold of your model to improve either sensitivity or specificity depending on how you intend to use the model.</p>\n<p>In a competition setting, knowing how trustful the labels are might help you go from a somehow working solution to a competitive one.</p>\n<p>In real life setting, which is the important setting, you do not want your model to mimic the current flaws and biases of the average doctor. What you want is the best possible model to detect cancer.</p>\n<p>I am not saying that it is easy or even possible to get a perfect labeling scheme. I am sure the hosts thought about this carefully and created the best possible dataset, nevertheless it's important to know the potential errors that can be hidden in the labeling process, so that maybe someone finds a way to efficiently deal with them. That's why I'd like to know the selection process for negative samples! Or as you say, how did they weed out the false negatives?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056107,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2022-12-05T18:28:57.263000",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> there's no way to completely eliminate false negatives, but in this case I would expect them to be a very minor issue. Taking the 20% false negative rate as a given, you might expect to see roughly 8,000 patients * 0.2 false negative rate *0.02 cancer rate = 32 false negatives. However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 2056609,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2022-12-06T09:37:27.537000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> ! Very clear, I guess we don't have to worry too much about false negative labels then!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056790,
          "author_name": "hrlmmm",
          "author_url": "",
          "post_date": "2022-12-06T12:36:38.020000",
          "content": "<p>Completely agreeing with you! This confusion about false negative cases has puzzled me for a long time. Glad to see this question and answer here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2057112,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-12-06T18:53:46.510000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> thanks for very nice write up.</p>\n<p>I am thinking if there is any way to use external data out there (i.e. vindr) which have BI-RADS in 1-5. <br>\nI am thinking of the following logic: <br>\nBI-RADS 4/5 (malignant) := BI-RADS 0 + biopsy<br>\nBI-RADS 2/3 (benign) := BI-RADS 2 or BI-RADS 0 + no-biopsy<br>\nBI-RADS 1 := BI-RADS 1 or BI-RADS is null</p>\n<p>Does this make sense at all?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2066075,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2022-12-15T11:40:44.533000",
          "content": "<p>I think it could be helpful, but most BI-RADS 4 are benign. There are actually further categories (4A, 4B, 4C) based on level of suspicion. Anything that has 2% or greater probability of cancer is recommended for biopsy.</p>\n<p>If the category 4 subtypes are provided, you could throw 4A into benign and 4C into malignant, ignoring 4B. Not perfect but would be better than just assuming 4 is malignant. </p>\n<p>On another note, I'm not sure if the VinDr-Mammo dataset (<a href=\"https://physionet.org/content/vindr-mammo/1.0.0/\" target=\"_blank\">https://physionet.org/content/vindr-mammo/1.0.0/</a>) is allowed. The data use agreement seems to specify that the data can only be used for research, which goes against the rules of the competition stating that there can be nothing limiting commercial use for the winning solutions. </p>\n<p>Maybe <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> can comment.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2147040,
      "author_name": "Jim Jing",
      "author_url": "",
      "post_date": "2023-02-16T11:23:49.753000",
      "content": "<p>Thanks for your introduction. Do you think it is reasonable to detect cancer instead abnormalities in Mammography ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2131870,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-02-06T13:00:19.070000",
      "content": "<p>Thank you for sharing this valuable post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2084389,
      "author_name": "zhenghuoer",
      "author_url": "",
      "post_date": "2023-01-03T13:09:19.047000",
      "content": "<p>Thanks for your intro, it's great useful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2057224,
      "author_name": "Saurav Solanki",
      "author_url": "",
      "post_date": "2022-12-06T21:32:08.363000",
      "content": "<p>It is worth reading and know the BTS of the process. This will definitely help us to build the model and understand the data well. <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2056551,
      "author_name": "Chibu",
      "author_url": "",
      "post_date": "2022-12-06T08:21:14.337000",
      "content": "<p>Thank you for this useful information, I didn't know about any of this… Will you be open to taking me on your team to learn from you, I have some knowledge in computer vision.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2052499,
      "author_name": "GarudA_Kai",
      "author_url": "",
      "post_date": "2022-12-02T08:56:12.917000",
      "content": "<p>Thanks for the information, but I have a question about the BI-RADS score. When I look at train.csv, I see that many patients do not have a BIRADS score. Among them are many patients who have cancer. Should I consider this as \"overlooked cancer\" or is there another reason, e.g. privacy?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2052521,
          "author_name": "hrlmmm",
          "author_url": "",
          "post_date": "2022-12-02T09:24:53.547000",
          "content": "<p>I think it's not a big due. Maybe some patients have no BIRADS score because they didn't record or share. Through the data we could find out,when the difficult_negetive_case is False and there is a BIRADS score,it must be BIRADS 0 for cancer 1, BIRADS 1 or 2 for cancer 0.If there isn't BIRADS score, We can assume that it follows the above rules.<br>\nIf you have any different opinions, I‘m glad to discuss it with you .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055394,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T03:57:14.533000",
          "content": "<p>site2 rarely reported birads.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2052353,
      "author_name": "explise",
      "author_url": "",
      "post_date": "2022-12-02T05:53:23.923000",
      "content": "<p>Thanks for sharing valuable information.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2052325,
      "author_name": "Arsive_AI",
      "author_url": "",
      "post_date": "2022-12-02T05:16:32.340000",
      "content": "<p>Informational. Thanks for explaining briefly</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2050575,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-11-30T21:04:09.860000",
      "content": "<p>This is incredibly helpful and clear. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2048910,
      "author_name": "Sixian Chen",
      "author_url": "",
      "post_date": "2022-11-29T19:16:29.097000",
      "content": "<p>Very informative!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2048638,
      "author_name": "Harbal Deep Sidhu",
      "author_url": "",
      "post_date": "2022-11-29T15:30:41.607000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for explaining mammography.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2049102,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-30T00:06:06.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2161461,
      "author_name": "Amit",
      "author_url": "",
      "post_date": "2023-02-27T14:14:44.283000",
      "content": "<p>Thank you for this useful info!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2151666,
      "author_name": "olafinho2",
      "author_url": "",
      "post_date": "2023-02-20T08:58:09.060000",
      "content": "<p>Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097281,
      "author_name": "Takayuki Higuchi",
      "author_url": "",
      "post_date": "2023-01-12T14:53:02.390000",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2085355,
      "author_name": "sho1_24",
      "author_url": "",
      "post_date": "2023-01-04T05:10:38.930000",
      "content": "<p>Thank you for useful information!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065928,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2022-12-15T07:56:26.427000",
      "content": "<p>Thanks for your great introduction. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2054933,
      "author_name": "Pavithra Devi M",
      "author_url": "",
      "post_date": "2022-12-04T15:10:15.503000",
      "content": "<p>thanks for sharing , very useful </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2051546,
      "author_name": "Tkikuchi",
      "author_url": "",
      "post_date": "2022-12-01T13:45:25.203000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2050451,
      "author_name": "Danial Monsefi Parapari",
      "author_url": "",
      "post_date": "2022-11-30T18:37:20.657000",
      "content": "<p>Thanks for the helpful details</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2050368,
      "author_name": "kiddu",
      "author_url": "",
      "post_date": "2022-11-30T17:35:52.647000",
      "content": "<p>Thanks for the valuable info mate!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2049354,
      "author_name": "zhenghuoer",
      "author_url": "",
      "post_date": "2022-11-30T04:31:26.693000",
      "content": "<p>Thanks for sharing~~~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2049334,
      "author_name": "hrlmmm",
      "author_url": "",
      "post_date": "2022-11-30T04:01:28.990000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2048515": "## What is mammography?\n\nMammography is the imaging modality of choice for breast cancer screening in women. It also plays an important role in the evaluation of other breast diseases in all patients. You can think of mammography as an X-ray for the breast. \n\nDifferent associations across the world have different guidelines for when to start and how often a patient should undergo breast cancer screening. The American College of Radiology and American College of Obstetrics and Gynecology recommend beginning **annual** screening at age 40. The United States Preventative Services Tasks Force recommends **biannual** screening at age 50. \n\nDepending on individual risk factors (BRCA gene mutations, strong family history, personal history of mantle radiation, etc.), screening may be recommended to begin earlier. Approximately 1 in 8 women will develop breast cancer in their lifetime. Screening allows for early detection and treatment. \n\n## What is a screening mammogram?\n\nA screening mammogram is the most common type of mammogram performed. During a screening exam, the patient goes to an imaging center for a scheduled appointment, and the mammogram technologist will perform two standard views: craniocaudal (CC) and mediolateral oblique (MLO). \n\nDuring mammography, the breast is put into compression, which can be quite uncomfortable for the patient! For the CC view, the breast is compressed from \"head to toe\" while for the MLO view, the breast is compressed \"side to side\" at an angle. This allows radiologists to localize a finding to a particular area in the breast and to determine what type of abnormality is present. The patient then leaves the imaging center after the appointment and awaits the results, which are required to be delivered within 30 days. \n\n## What is BI-RADS?\n\nThe Breast Imaging Reporting and Data System \"provides standardized breast imaging terminology, report organization, assessment structure and a classification system for mammography, ultrasound and MRI of the breast\"(https://www.acr.org/Clinical-Resources/Reporting-and-Data-Systems/Bi-Rads). It is a standard lexicon used by radiologists that allows for improved peer review and quality assurance mechanisms with the goal of improving patient care. \n\nBI-RADS is not only used for mammography but also for breast ultrasound and breast MRI. It essentially defines the vocabulary for breast imaging, a summary of which can be found here: https://www.acr.org/-/media/ACR/Files/RADS/BI-RADS/BIRADS-Reference-Card.pdf\n\nFor the purposes of this challenge, the most important element of BI-RADS to understand is the **BI-RADS score**.\n\nThere are 7 main BI-RADS scores, or categories:\n0 - Need additional imaging evaluation\n1 - Negative\n2 - Benign\n3 - Probably Benign\n4 - Suspicious\n5 - Highly Suggestive of Malignancy\n6 - Known Biopsy-Proven Malignancy\n\n**Screening mammograms can only be assigned a BI-RADS score of 0, 1, or 2.** This is why the dataset only contains 3 of the BI-RADS categories. The reason why will be clear as we talk about the breast imaging workflow.\n\n## What is the breast imaging workflow?\n\nWe talked briefly about screening mammography earlier in this post. If a patient's screening mammogram is given BI-RADS 1 or 2, then they will be due for another screening at whatever interval they and their physician decide. \n\nIf a screening mammogram is given BI-RADS 0, this means the interpreting radiologist saw an abnormality (mass, calcification, asymmetry) that warranted further evaluation. **You cannot diagnose cancer on a screening mammogram.** The patient is then \"called back\" for additional imaging. \n\nAt the next appointment, the patient will undergo a **diagnostic** mammogram. The standard MLO and CC views are once obtained, but typically these exams are performed **online**, which means the patient does not leave until all necessary imaging is obtained. After the standard views are obtained, the radiologist will look at the images and determine if additional, more specific, views are needed. Examples include magnification views to examine calcifications or \"spot compression\" views to see if a suspicious area of tissue goes away if more compression is applied. \n\nOftentimes, a breast ultrasound is performed as well for additional evaluation. When the radiologist is satisfied with all the imaging, they will assign a final BI-RADS score, which can vary from 1 to 6. \n\nBI-RADS 3 indicates a probably benign finding that will be followed for a period of time (usually 6, 12, and 24 months). This means the patient returns for a **diagnostic** appointment at each interval, and the interval is more frequent than regular screening. \n\nBI-RADS 4 indicates a suspicious finding and is further subdivided into 4A, 4B, and 4C, depending on the level of suspicion that the finding may represent cancer. **BI-RADS 4 means you are recommending a biopsy for tissue diagnosis.** \n\nBI-RADS 5 indicates a highly suspicious finding- a greater than 95% chance that you believe the finding is cancer. BI-RADS 5 also means you are recommending a biopsy/tissue sampling. In fact, if you are assigning BI-RADS 5, even if the biopsy is negative, you would suspect that something went wrong during the biopsy and recommend repeat biopsy or even surgical excision because you believe the finding to be that suspicious.\n\nBI-RADS 6 is given when the patient has a known malignancy. These are usually pre-treatment planning cases or cases where the patient may be undergoing chemotherapy for a cancer that cannot be surgically treated. \n\n## Our task\n\nFor this task, we are focused on **screening mammograms**, which again means that only scores of **BI-RADS 0, 1, or 2** are possible. One can essentially think of 0 as \"abnormal\" and 1 and 2 as \"normal.\" There can be subjectivity in assigning BI-RADS 1 or 2. For example, if there are stable findings that are almost certainly benign breast cysts, the mammogram may be assigned 1 or 2 depending on the radiologist. \n\nHowever, our task in the challenge is to predict **cancer or no cancer**, a binary value. This may be obvious, but not all BI-RADS 0 cases have cancer. The final cancer label will depend on the outcome of the diagnostic mammogram and the biopsy results, if obtained. \n\nIn this post, I wanted to give competitors a sense of what mammography is and how the breast imaging workflow is designed so they can contextualize the challenge task within the broader scope of breast cancer diagnosis and treatment. There is a lot more to mammography, but I hope people find this introduction helpful!",
    "2052881": "@vaillant thank you for this useful information.\n\nSo if someone gets a BI-RADS of 1 or 2, there will not be any additional actions in the near future. This means that no biopsy will be performed and so there is no chance that the images will be labeled as cancer.\n\nDo you know, in general, what is the False Negative rate for patients with BI-RADS 1 and 2 for an average physician? Is it around 0.01%, 0.1%, 1%? ",
    "2057112": "@vaillant thanks for very nice write up.\n\nI am thinking if there is any way to use external data out there (i.e. vindr) which have BI-RADS in 1-5. \nI am thinking of the following logic: \nBI-RADS 4/5 (malignant) := BI-RADS 0 + biopsy\nBI-RADS 2/3 (benign) := BI-RADS 2 or BI-RADS 0 + no-biopsy\nBI-RADS 1 := BI-RADS 1 or BI-RADS is null\n\nDoes this make sense at all?",
    "2147040": "Thanks for your introduction. Do you think it is reasonable to detect cancer instead abnormalities in Mammography ? ",
    "2131870": "Thank you for sharing this valuable post!",
    "2084389": "Thanks for your intro, it's great useful!",
    "2057224": "It is worth reading and know the BTS of the process. This will definitely help us to build the model and understand the data well. @vaillant ",
    "2056551": "Thank you for this useful information, I didn't know about any of this... Will you be open to taking me on your team to learn from you, I have some knowledge in computer vision.",
    "2052499": "Thanks for the information, but I have a question about the BI-RADS score. When I look at train.csv, I see that many patients do not have a BIRADS score. Among them are many patients who have cancer. Should I consider this as \"overlooked cancer\" or is there another reason, e.g. privacy?",
    "2052353": "Thanks for sharing valuable information.",
    "2052325": "Informational. Thanks for explaining briefly",
    "2050575": "This is incredibly helpful and clear. Thank you!",
    "2048910": "Very informative!",
    "2048638": "Thanks @vaillant for explaining mammography.",
    "2049102": "",
    "2161461": "Thank you for this useful info!",
    "2151666": "Thank you!",
    "2097281": "Thanks for sharing!!",
    "2085355": "Thank you for useful information!",
    "2065928": "Thanks for your great introduction. ",
    "2054933": "thanks for sharing , very useful ",
    "2051546": "Thanks for sharing!",
    "2050451": "Thanks for the helpful details",
    "2050368": "Thanks for the valuable info mate!",
    "2049354": "Thanks for sharing~~~",
    "2049334": "Thanks for sharing"
  }
}