{
  "id": 373208,
  "title": "Weird Mammograms in the Dataset",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/373208",
  "author_name": "outwrest",
  "post_date": "2022-12-20T07:40:14.306000",
  "votes": 19,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've noticed that patient_id <code>27770</code> has some weird mammograms. I believe that this is an error due to the vast amount of noise (and no visible tissue) and it stands out against all the other images. I made a <a href=\"https://www.kaggle.com/code/outwrest/weird-mammograms\" target=\"_blank\">quick notebook</a> to try out other processing methods, maybe I am doing something wrong. Only 4 images are affected and below are a couple of examples.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F815c80a278576e43b2a194ceaa7b2856%2F__results___3_1.png?generation=1671521622197429&amp;alt=media\" alt=\"Noisy Mammogram\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F6ee400331038ac27076fbc3f2849ac87%2F__results___3_0.png?generation=1671521642714482&amp;alt=media\" alt=\"Noisy Mammogram\"></p>\n<p>I've looked into public preprocessed datasets and I see the same noisy image. What do you guys think?</p>",
  "messages": [
    {
      "id": 2070602,
      "postDate": "2022-12-20T07:40:14.307Z",
      "content": "<p>I've noticed that patient_id <code>27770</code> has some weird mammograms. I believe that this is an error due to the vast amount of noise (and no visible tissue) and it stands out against all the other images. I made a <a href=\"https://www.kaggle.com/code/outwrest/weird-mammograms\" target=\"_blank\">quick notebook</a> to try out other processing methods, maybe I am doing something wrong. Only 4 images are affected and below are a couple of examples.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F815c80a278576e43b2a194ceaa7b2856%2F__results___3_1.png?generation=1671521622197429&amp;alt=media\" alt=\"Noisy Mammogram\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F6ee400331038ac27076fbc3f2849ac87%2F__results___3_0.png?generation=1671521642714482&amp;alt=media\" alt=\"Noisy Mammogram\"></p>\n<p>I've looked into public preprocessed datasets and I see the same noisy image. What do you guys think?</p>",
      "rawMarkdown": "I've noticed that patient_id `27770` has some weird mammograms. I believe that this is an error due to the vast amount of noise (and no visible tissue) and it stands out against all the other images. I made a [quick notebook](https://www.kaggle.com/code/outwrest/weird-mammograms) to try out other processing methods, maybe I am doing something wrong. Only 4 images are affected and below are a couple of examples.\n\n![Noisy Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F815c80a278576e43b2a194ceaa7b2856%2F__results___3_1.png?generation=1671521622197429&alt=media)\n\n![Noisy Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F6ee400331038ac27076fbc3f2849ac87%2F__results___3_0.png?generation=1671521642714482&alt=media)\n\nI've looked into public preprocessed datasets and I see the same noisy image. What do you guys think?",
      "votes": 19
    },
    {
      "id": 2072051,
      "postDate": "2022-12-21T16:33:56.460Z",
      "content": "<p>Hi, I also found patient id 822 picture id 1942326353 is just a white picture after changing to PNG.<br>\nI checked many datasets and all the same.</p>",
      "rawMarkdown": "Hi, I also found patient id 822 picture id 1942326353 is just a white picture after changing to PNG.\nI checked many datasets and all the same.",
      "votes": 2,
      "replies": [
        {
          "id": 2072155,
          "postDate": "2022-12-21T19:07:35.197Z",
          "content": "<p>Good catch, thanks! I added it notebook.</p>",
          "rawMarkdown": "Good catch, thanks! I added it notebook."
        }
      ]
    },
    {
      "id": 2071592,
      "postDate": "2022-12-21T07:54:57.040Z",
      "content": "<p>I updated the notebook to include patient_id <code>45629</code> because it seems like it was <em>weirdly</em> annotated and centered (only 2 are annotated, 4 are centered, and 8 mammograms total for the patient_id.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fda50d51a74927ffb71c3b160a2b735e1%2F__results___3_9.png?generation=1671606371435285&amp;alt=media\" alt=\"Weird Mammogram\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F838ba46b329bdb2710cf6f12c97be857%2F__results___3_7.png?generation=1671606388705875&amp;alt=media\" alt=\"Weird Mammogram\"></p>\n<ul>\n<li>The first one has a circle on an area of the breast.</li>\n<li>The next image has the same circle but with the annotation <code>mag cc, mag ml, full 90 tomo combo</code></li>\n</ul>\n<p>Does the circle mean anything? (It is just there, but also it is right in the middle?) I also asked ChatGPT about the annotation since I was unsure about the meaning. It turns out it refers to the type of imaging that was done. I believe that they're both meaningless</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F74d850ff150036b50ddbd30878332c4b%2FCapture.PNG?generation=1671607408043463&amp;alt=media\" alt=\"W ChatGPT\"></p>\n<p>While not specific to this post, here are some things to consider when building a robust preprocessing system, from my analysis of ROI extraction. I hope it may be useful for other competitors. </p>\n<ul>\n<li>Be cautious with the dataset, it is not fully consistent. While most are generally great, there's a minority that--I believe--can impact model performance. I showed some extreme examples here but most of the inconsistencies are way more subtle. </li>\n<li>Always try to find outliers and work toward addressing them. I kept track of different heuristics about location, position, size, etc., and classified outliers.</li>\n</ul>\n<p>I invite everyone to check out the <a href=\"https://cs.nyu.edu/~kgeras/reports/datav1.0.pdf\" target=\"_blank\">The NYU Breast Cancer Screening Dataset</a> paper. The dataset is private but they go into detail about how they processed over <strong>1 million</strong> mammography images (I believe there are some implementations in the NYU GitHub repo if you are interested).</p>\n<p>I think we are only scratching the surface and a lot more to do, let me know what you guys think!</p>",
      "rawMarkdown": "I updated the notebook to include patient_id `45629` because it seems like it was *weirdly* annotated and centered (only 2 are annotated, 4 are centered, and 8 mammograms total for the patient_id.\n\n![Weird Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fda50d51a74927ffb71c3b160a2b735e1%2F__results___3_9.png?generation=1671606371435285&alt=media)\n\n![Weird Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F838ba46b329bdb2710cf6f12c97be857%2F__results___3_7.png?generation=1671606388705875&alt=media)\n\n* The first one has a circle on an area of the breast.\n* The next image has the same circle but with the annotation `mag cc, mag ml, full 90 tomo combo`\n\nDoes the circle mean anything? (It is just there, but also it is right in the middle?) I also asked ChatGPT about the annotation since I was unsure about the meaning. It turns out it refers to the type of imaging that was done. I believe that they're both meaningless\n\n![W ChatGPT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F74d850ff150036b50ddbd30878332c4b%2FCapture.PNG?generation=1671607408043463&alt=media)\n\nWhile not specific to this post, here are some things to consider when building a robust preprocessing system, from my analysis of ROI extraction. I hope it may be useful for other competitors. \n\n* Be cautious with the dataset, it is not fully consistent. While most are generally great, there's a minority that--I believe--can impact model performance. I showed some extreme examples here but most of the inconsistencies are way more subtle. \n* Always try to find outliers and work toward addressing them. I kept track of different heuristics about location, position, size, etc., and classified outliers.\n\nI invite everyone to check out the [The NYU Breast Cancer Screening Dataset](https://cs.nyu.edu/~kgeras/reports/datav1.0.pdf) paper. The dataset is private but they go into detail about how they processed over **1 million** mammography images (I believe there are some implementations in the NYU GitHub repo if you are interested).\n\nI think we are only scratching the surface and a lot more to do, let me know what you guys think!",
      "votes": 2
    },
    {
      "id": 2075808,
      "postDate": "2022-12-25T22:41:36.003Z",
      "content": "<p>Has anyone noticed decoding failures with dicomsdl when multithreading? </p>\n<p>edit: (16+ threads, can't reproduce)</p>",
      "rawMarkdown": "Has anyone noticed decoding failures with dicomsdl when multithreading? \n\nedit: (16+ threads, can't reproduce)"
    },
    {
      "id": 2072721,
      "postDate": "2022-12-22T11:18:56.547Z",
      "content": "<p>As I wrote on yours Notebook: interesting, intriguing approach.</p>",
      "rawMarkdown": "As I wrote on yours Notebook: interesting, intriguing approach."
    },
    {
      "id": 2071337,
      "postDate": "2022-12-20T22:21:32.420Z",
      "content": "<p>All four of the original DICOM images of patient_id 27770 are just noise. It looks like an error during conversion to JPG transfer syntax. I don't think you will be able to extract anything useful from them.</p>",
      "rawMarkdown": "All four of the original DICOM images of patient_id 27770 are just noise. It looks like an error during conversion to JPG transfer syntax. I don't think you will be able to extract anything useful from them."
    },
    {
      "id": 2071135,
      "postDate": "2022-12-20T17:21:17.513Z",
      "content": "<p>Hey just a heads up to anyone doing preprocessing, I also found an image where background is noisy and not all zeros (especially if you’re doing breast ROI extraction). There seems to be a few data points which are inconsistent in this dataset </p>",
      "rawMarkdown": "Hey just a heads up to anyone doing preprocessing, I also found an image where background is noisy and not all zeros (especially if you’re doing breast ROI extraction). There seems to be a few data points which are inconsistent in this dataset "
    }
  ],
  "comments": [
    {
      "id": 2072051,
      "author_name": "nicehzj",
      "author_url": "",
      "post_date": "2022-12-21T16:33:56.460000",
      "content": "<p>Hi, I also found patient id 822 picture id 1942326353 is just a white picture after changing to PNG.<br>\nI checked many datasets and all the same.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2072155,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "2022-12-21T19:07:35.197000",
          "content": "<p>Good catch, thanks! I added it notebook.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2071592,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-12-21T07:54:57.040000",
      "content": "<p>I updated the notebook to include patient_id <code>45629</code> because it seems like it was <em>weirdly</em> annotated and centered (only 2 are annotated, 4 are centered, and 8 mammograms total for the patient_id.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fda50d51a74927ffb71c3b160a2b735e1%2F__results___3_9.png?generation=1671606371435285&amp;alt=media\" alt=\"Weird Mammogram\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F838ba46b329bdb2710cf6f12c97be857%2F__results___3_7.png?generation=1671606388705875&amp;alt=media\" alt=\"Weird Mammogram\"></p>\n<ul>\n<li>The first one has a circle on an area of the breast.</li>\n<li>The next image has the same circle but with the annotation <code>mag cc, mag ml, full 90 tomo combo</code></li>\n</ul>\n<p>Does the circle mean anything? (It is just there, but also it is right in the middle?) I also asked ChatGPT about the annotation since I was unsure about the meaning. It turns out it refers to the type of imaging that was done. I believe that they're both meaningless</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F74d850ff150036b50ddbd30878332c4b%2FCapture.PNG?generation=1671607408043463&amp;alt=media\" alt=\"W ChatGPT\"></p>\n<p>While not specific to this post, here are some things to consider when building a robust preprocessing system, from my analysis of ROI extraction. I hope it may be useful for other competitors. </p>\n<ul>\n<li>Be cautious with the dataset, it is not fully consistent. While most are generally great, there's a minority that--I believe--can impact model performance. I showed some extreme examples here but most of the inconsistencies are way more subtle. </li>\n<li>Always try to find outliers and work toward addressing them. I kept track of different heuristics about location, position, size, etc., and classified outliers.</li>\n</ul>\n<p>I invite everyone to check out the <a href=\"https://cs.nyu.edu/~kgeras/reports/datav1.0.pdf\" target=\"_blank\">The NYU Breast Cancer Screening Dataset</a> paper. The dataset is private but they go into detail about how they processed over <strong>1 million</strong> mammography images (I believe there are some implementations in the NYU GitHub repo if you are interested).</p>\n<p>I think we are only scratching the surface and a lot more to do, let me know what you guys think!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2075808,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-12-25T22:41:36.003000",
      "content": "<p>Has anyone noticed decoding failures with dicomsdl when multithreading? </p>\n<p>edit: (16+ threads, can't reproduce)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2072721,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-12-22T11:18:56.547000",
      "content": "<p>As I wrote on yours Notebook: interesting, intriguing approach.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2071337,
      "author_name": "David Roberts",
      "author_url": "",
      "post_date": "2022-12-20T22:21:32.420000",
      "content": "<p>All four of the original DICOM images of patient_id 27770 are just noise. It looks like an error during conversion to JPG transfer syntax. I don't think you will be able to extract anything useful from them.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2071135,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-12-20T17:21:17.513000",
      "content": "<p>Hey just a heads up to anyone doing preprocessing, I also found an image where background is noisy and not all zeros (especially if you’re doing breast ROI extraction). There seems to be a few data points which are inconsistent in this dataset </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2070602": "I've noticed that patient_id `27770` has some weird mammograms. I believe that this is an error due to the vast amount of noise (and no visible tissue) and it stands out against all the other images. I made a [quick notebook](https://www.kaggle.com/code/outwrest/weird-mammograms) to try out other processing methods, maybe I am doing something wrong. Only 4 images are affected and below are a couple of examples.\n\n![Noisy Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F815c80a278576e43b2a194ceaa7b2856%2F__results___3_1.png?generation=1671521622197429&alt=media)\n\n![Noisy Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F6ee400331038ac27076fbc3f2849ac87%2F__results___3_0.png?generation=1671521642714482&alt=media)\n\nI've looked into public preprocessed datasets and I see the same noisy image. What do you guys think?",
    "2072051": "Hi, I also found patient id 822 picture id 1942326353 is just a white picture after changing to PNG.\nI checked many datasets and all the same.",
    "2071592": "I updated the notebook to include patient_id `45629` because it seems like it was *weirdly* annotated and centered (only 2 are annotated, 4 are centered, and 8 mammograms total for the patient_id.\n\n![Weird Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2Fda50d51a74927ffb71c3b160a2b735e1%2F__results___3_9.png?generation=1671606371435285&alt=media)\n\n![Weird Mammogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F838ba46b329bdb2710cf6f12c97be857%2F__results___3_7.png?generation=1671606388705875&alt=media)\n\n* The first one has a circle on an area of the breast.\n* The next image has the same circle but with the annotation `mag cc, mag ml, full 90 tomo combo`\n\nDoes the circle mean anything? (It is just there, but also it is right in the middle?) I also asked ChatGPT about the annotation since I was unsure about the meaning. It turns out it refers to the type of imaging that was done. I believe that they're both meaningless\n\n![W ChatGPT](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5918909%2F74d850ff150036b50ddbd30878332c4b%2FCapture.PNG?generation=1671607408043463&alt=media)\n\nWhile not specific to this post, here are some things to consider when building a robust preprocessing system, from my analysis of ROI extraction. I hope it may be useful for other competitors. \n\n* Be cautious with the dataset, it is not fully consistent. While most are generally great, there's a minority that--I believe--can impact model performance. I showed some extreme examples here but most of the inconsistencies are way more subtle. \n* Always try to find outliers and work toward addressing them. I kept track of different heuristics about location, position, size, etc., and classified outliers.\n\nI invite everyone to check out the [The NYU Breast Cancer Screening Dataset](https://cs.nyu.edu/~kgeras/reports/datav1.0.pdf) paper. The dataset is private but they go into detail about how they processed over **1 million** mammography images (I believe there are some implementations in the NYU GitHub repo if you are interested).\n\nI think we are only scratching the surface and a lot more to do, let me know what you guys think!",
    "2075808": "Has anyone noticed decoding failures with dicomsdl when multithreading? \n\nedit: (16+ threads, can't reproduce)",
    "2072721": "As I wrote on yours Notebook: interesting, intriguing approach.",
    "2071337": "All four of the original DICOM images of patient_id 27770 are just noise. It looks like an error during conversion to JPG transfer syntax. I don't think you will be able to extract anything useful from them.",
    "2071135": "Hey just a heads up to anyone doing preprocessing, I also found an image where background is noisy and not all zeros (especially if you’re doing breast ROI extraction). There seems to be a few data points which are inconsistent in this dataset "
  }
}