{
  "id": 372186,
  "title": "Classfication to Object Detection [GradCAM -> BBox]",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372186",
  "author_name": "Awsaf",
  "post_date": "2022-12-14T16:20:23.109000",
  "votes": 36,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I am wondering if we can convert this <strong>cancer detection</strong> problem from <code>classification</code> to <code>object detection</code>. There are couple of good reasons for considering object detection, such as cancer pixels (object) contains very small portion of the image. If we consider full dicom (w/o roi) then it becomes even worse. The model has to overcome a lot of distractions due to such imbalance. In object detection, model can optimize more easily using <strong>bounding box</strong>. Then there are some models like <strong>yolov5</strong>, <strong>yolov7</strong> and so on..</p>\n<p>But sadly we don't have any <strong>bounding box</strong>. But we do have access to <strong>class activation map (Grad-CAM)</strong> which shows parts of an input image that most impact the classification score. What if we convert the  <strong>grad-cam</strong> to <strong>bounding box (bbox)</strong> ? I published following notebook to carry out this experiment with <strong>WandB</strong>.</p>\n<h2>Notebook:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-gradcam-to-bbox\" target=\"_blank\">RSNA-BCD: GradCAM to BBox</a></li>\n<li>RSNA Breast Cancer Detection [YOLOv5] (Coming Soon…)</li>\n</ul>\n<h2>Steps:</h2>\n<ol>\n<li>Take <strong>cancer</strong> images only.</li>\n<li>Apply filter to <strong>cancer</strong> images to take only images with <strong>good confidence</strong>. The threshold  for <code>confidence</code> can be obtained by maximizing <code>pF1</code> score like we're doing in classification.</li>\n<li>Convert <strong>gradcam</strong> to <strong>mask</strong> using <code>thresholding</code> then obtain <strong>bbox</strong> from <strong>mask</strong> using <code>largest contour</code>, so in a nutshell, image -&gt; gradcam -&gt; mask -&gt; bbox</li>\n<li>In training, use only the <strong>filtered cancer</strong> images obtained from the above steps.</li>\n<li>Apply pseudo labeling on other <strong>cancer</strong> images to obtain their labels.</li>\n</ol>\n<h2>WandB</h2>\n<p><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb96243b27b03d4635f6a9242b2c29372%2Fgradcam2bbox-wandb.PNG?generation=1671034084366419&amp;alt=media\" alt=\"\"></p>\n<h2>Samples</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F31e9e3bd9915e6fc120c79f212d366e2%2Fgradcam2bbox.jpg?generation=1671033517214365&amp;alt=media\"></p>\n<h2>Key-Points</h2>\n<ul>\n<li>GradCAM doesn't always point to the right location even when the prediction is correct. So, we need to use this method very carefully. Usually, high confidence means grad-cam is more likely to point to the correct location.</li>\n<li>Better classification models definitely generate better bounding box, hence need accurate models.</li>\n<li>We can ensemble different mdoels' gradcams by using method like <strong>Weighted Boxes Fusion (WBF)</strong> by considering model's prediction as confidence score of <strong>bounding box</strong>.</li>\n</ul>",
  "messages": [
    {
      "id": 2065413,
      "postDate": "2022-12-14T16:20:23.110Z",
      "content": "<p>I am wondering if we can convert this <strong>cancer detection</strong> problem from <code>classification</code> to <code>object detection</code>. There are couple of good reasons for considering object detection, such as cancer pixels (object) contains very small portion of the image. If we consider full dicom (w/o roi) then it becomes even worse. The model has to overcome a lot of distractions due to such imbalance. In object detection, model can optimize more easily using <strong>bounding box</strong>. Then there are some models like <strong>yolov5</strong>, <strong>yolov7</strong> and so on..</p>\n<p>But sadly we don't have any <strong>bounding box</strong>. But we do have access to <strong>class activation map (Grad-CAM)</strong> which shows parts of an input image that most impact the classification score. What if we convert the  <strong>grad-cam</strong> to <strong>bounding box (bbox)</strong> ? I published following notebook to carry out this experiment with <strong>WandB</strong>.</p>\n<h2>Notebook:</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-gradcam-to-bbox\" target=\"_blank\">RSNA-BCD: GradCAM to BBox</a></li>\n<li>RSNA Breast Cancer Detection [YOLOv5] (Coming Soon…)</li>\n</ul>\n<h2>Steps:</h2>\n<ol>\n<li>Take <strong>cancer</strong> images only.</li>\n<li>Apply filter to <strong>cancer</strong> images to take only images with <strong>good confidence</strong>. The threshold  for <code>confidence</code> can be obtained by maximizing <code>pF1</code> score like we're doing in classification.</li>\n<li>Convert <strong>gradcam</strong> to <strong>mask</strong> using <code>thresholding</code> then obtain <strong>bbox</strong> from <strong>mask</strong> using <code>largest contour</code>, so in a nutshell, image -&gt; gradcam -&gt; mask -&gt; bbox</li>\n<li>In training, use only the <strong>filtered cancer</strong> images obtained from the above steps.</li>\n<li>Apply pseudo labeling on other <strong>cancer</strong> images to obtain their labels.</li>\n</ol>\n<h2>WandB</h2>\n<p><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb96243b27b03d4635f6a9242b2c29372%2Fgradcam2bbox-wandb.PNG?generation=1671034084366419&amp;alt=media\" alt=\"\"></p>\n<h2>Samples</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F31e9e3bd9915e6fc120c79f212d366e2%2Fgradcam2bbox.jpg?generation=1671033517214365&amp;alt=media\"></p>\n<h2>Key-Points</h2>\n<ul>\n<li>GradCAM doesn't always point to the right location even when the prediction is correct. So, we need to use this method very carefully. Usually, high confidence means grad-cam is more likely to point to the correct location.</li>\n<li>Better classification models definitely generate better bounding box, hence need accurate models.</li>\n<li>We can ensemble different mdoels' gradcams by using method like <strong>Weighted Boxes Fusion (WBF)</strong> by considering model's prediction as confidence score of <strong>bounding box</strong>.</li>\n</ul>",
      "rawMarkdown": "I am wondering if we can convert this **cancer detection** problem from `classification` to `object detection`. There are couple of good reasons for considering object detection, such as cancer pixels (object) contains very small portion of the image. If we consider full dicom (w/o roi) then it becomes even worse. The model has to overcome a lot of distractions due to such imbalance. In object detection, model can optimize more easily using **bounding box**. Then there are some models like **yolov5**, **yolov7** and so on..\n\nBut sadly we don't have any **bounding box**. But we do have access to **class activation map (Grad-CAM)** which shows parts of an input image that most impact the classification score. What if we convert the  **grad-cam** to **bounding box (bbox)** ? I published following notebook to carry out this experiment with **WandB**.\n\n## Notebook:\n* [RSNA-BCD: GradCAM to BBox](https://www.kaggle.com/code/awsaf49/rsna-bcd-gradcam-to-bbox)\n* RSNA Breast Cancer Detection [YOLOv5] (Coming Soon...)\n\n## Steps:\n1. Take **cancer** images only.\n2. Apply filter to **cancer** images to take only images with **good confidence**. The threshold  for `confidence` can be obtained by maximizing `pF1` score like we're doing in classification.\n4. Convert **gradcam** to **mask** using `thresholding` then obtain **bbox** from **mask** using `largest contour`, so in a nutshell, image -> gradcam -> mask -> bbox\n3. In training, use only the **filtered cancer** images obtained from the above steps.\n4. Apply pseudo labeling on other **cancer** images to obtain their labels.\n\n## WandB\n<span style=\"color: #000508; font-family: Segoe UI; font-size: 1.5em; font-weight: 300;\"><a href=\"https://wandb.ai/awsaf49/rsna-bcd-gradcam2bbox\">View the Complete Dashboard Here ⮕</a></span>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb96243b27b03d4635f6a9242b2c29372%2Fgradcam2bbox-wandb.PNG?generation=1671034084366419&alt=media)\n\n## Samples\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F31e9e3bd9915e6fc120c79f212d366e2%2Fgradcam2bbox.jpg?generation=1671033517214365&alt=media\" height=\"900\">\n\n## Key-Points\n* GradCAM doesn't always point to the right location even when the prediction is correct. So, we need to use this method very carefully. Usually, high confidence means grad-cam is more likely to point to the correct location.\n* Better classification models definitely generate better bounding box, hence need accurate models.\n* We can ensemble different mdoels' gradcams by using method like **Weighted Boxes Fusion (WBF)** by considering model's prediction as confidence score of **bounding box**.",
      "votes": 36
    },
    {
      "id": 2067791,
      "postDate": "2022-12-17T06:48:46.457Z",
      "content": "<p><img src=\"https://i.ibb.co/f8YBSgd/mia-structure.png\" alt=\"https://i.ibb.co/f8YBSgd/mia-structure.png\"><br>\n<a href=\"https://github.com/nyukat/GMIC\" target=\"_blank\">https://github.com/nyukat/GMIC</a></p>",
      "rawMarkdown": "![https://i.ibb.co/f8YBSgd/mia-structure.png](https://i.ibb.co/f8YBSgd/mia-structure.png)\nhttps://github.com/nyukat/GMIC",
      "votes": 3
    },
    {
      "id": 2065432,
      "postDate": "2022-12-14T16:49:00.170Z",
      "content": "<p>a better solution is:</p>\n<ol>\n<li>find open source dataset with lesion annotation (there are many of them, some are box annotations and some are mask annotations) i think you can get about 3k to 4k annotations.</li>\n<li>train a segmentation model using those external dataset</li>\n<li>apply the segmentation model on kaggle dataset to make output predicted probability as a new channel </li>\n<li>now you have input=gray+probability to train a new kaggle model.</li>\n</ol>\n<p>shortcut:</p>\n<ul>\n<li>use opensource segmentation model and avoid step.2 for a quick test.<br>\ne.g. <a href=\"https://github.com/nyukat/\" target=\"_blank\">https://github.com/nyukat/</a></li>\n</ul>\n<p>these models are trained on millions of mammography images.<br>\nthe approach described above is also based on their papers on the github</p>",
      "rawMarkdown": "a better solution is:\n1.  find open source dataset with lesion annotation (there are many of them, some are box annotations and some are mask annotations) i think you can get about 3k to 4k annotations.\n2. train a segmentation model using those external dataset\n3. apply the segmentation model on kaggle dataset to make output predicted probability as a new channel \n4. now you have input=gray+probability to train a new kaggle model.\n\nshortcut:\n- use opensource segmentation model and avoid step.2 for a quick test.\ne.g. https://github.com/nyukat/\n\nthese models are trained on millions of mammography images.\nthe approach described above is also based on their papers on the github",
      "votes": 4,
      "replies": [
        {
          "id": 2065440,
          "postDate": "2022-12-14T16:53:57.733Z",
          "content": "<p>just curious  …<br>\n<img src=\"https://i.ibb.co/vq5F86x/Selection-196.png\" alt=\"https://i.ibb.co/vq5F86x/Selection-196.png\"></p>\n<p>actually, GPTchat is good at answering general machine-learning problems. i have asked other questions regarding class imbalance, etc, and GPTchat gives good answers … maybe because he is just talking about his own personal experiences, how he is trained by openAI</p>",
          "rawMarkdown": "just curious  ...\n![https://i.ibb.co/vq5F86x/Selection-196.png](https://i.ibb.co/vq5F86x/Selection-196.png)\n\n\nactually, GPTchat is good at answering general machine-learning problems. i have asked other questions regarding class imbalance, etc, and GPTchat gives good answers ... maybe because he is just talking about his own personal experiences, how he is trained by openAI",
          "votes": 7,
          "replies": [
            {
              "id": 2067762,
              "postDate": "2022-12-17T06:00:08.557Z",
              "content": "<p>Did GPTChat write the code for you? </p>",
              "rawMarkdown": "Did GPTChat write the code for you? \n"
            }
          ]
        }
      ]
    },
    {
      "id": 2070287,
      "postDate": "2022-12-19T19:57:03.957Z",
      "content": "<p>I created a yolov5 POC here .. for the test I used the same images as the grad cam notebook listed above.</p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fd174d1dfa9098c927d7b8eb95f422c04%2FScreenshot%202022-12-19%2011.54.30%20AM.png?generation=1671479701605290&amp;alt=media\" alt=\"\"></p>\n<p>It's using the yolov5 models trained here - <a href=\"https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0\" target=\"_blank\">https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0</a></p>\n<p>I reduced the conf thresh to 0.1 .. other params can be tweaked.  Overlapping boundary boxes will probably want to be merged when pushing them through the next stage of the model (efnet, resnet, etc that detects benign/cancer).  Likely the yolo models need more training as well, but it's a start.</p>",
      "rawMarkdown": "I created a yolov5 POC here .. for the test I used the same images as the grad cam notebook listed above.\n\nhttps://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fd174d1dfa9098c927d7b8eb95f422c04%2FScreenshot%202022-12-19%2011.54.30%20AM.png?generation=1671479701605290&alt=media)\n\nIt's using the yolov5 models trained here - https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0\n\nI reduced the conf thresh to 0.1 .. other params can be tweaked.  Overlapping boundary boxes will probably want to be merged when pushing them through the next stage of the model (efnet, resnet, etc that detects benign/cancer).  Likely the yolo models need more training as well, but it's a start.",
      "votes": 1,
      "replies": [
        {
          "id": 2070457,
          "postDate": "2022-12-20T02:54:17.333Z",
          "content": "<p>That's great =) Try pseudo labeling to generate labels for restrest of the positive images. Then re-train.</p>",
          "rawMarkdown": "That's great =) Try pseudo labeling to generate labels for restrest of the positive images. Then re-train.",
          "votes": 1,
          "replies": [
            {
              "id": 2070471,
              "postDate": "2022-12-20T03:32:06.990Z",
              "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> </p>\n<p>I think the right idea is to go hard on </p>\n<p>Pham, H. H., Nguyen Trung, H., &amp; Nguyen, H. Q. (2022). VinDr-Mammo: A large-scale benchmark dataset for computer-aided detection and diagnosis in full-field digital mammography (version 1.0.0). PhysioNet. <a href=\"https://doi.org/10.13026/br2v-7517\" target=\"_blank\">https://doi.org/10.13026/br2v-7517</a>.</p>\n<p>It's FFDM and it's huge (5000 patients), and it has finding_annotations.csv to train yolo.</p>\n<p>But for folks competing, they might not be able to. I've reached out to the physionet folks to see if it would be permissible.</p>",
              "rawMarkdown": "@awsaf49 \n\nI think the right idea is to go hard on \n\nPham, H. H., Nguyen Trung, H., & Nguyen, H. Q. (2022). VinDr-Mammo: A large-scale benchmark dataset for computer-aided detection and diagnosis in full-field digital mammography (version 1.0.0). PhysioNet. https://doi.org/10.13026/br2v-7517.\n\nIt's FFDM and it's huge (5000 patients), and it has finding_annotations.csv to train yolo.\n\nBut for folks competing, they might not be able to. I've reached out to the physionet folks to see if it would be permissible."
            },
            {
              "id": 2070483,
              "postDate": "2022-12-20T04:00:03.897Z",
              "content": "<p>Do they annotate Mass or Cancer? Are they same thing?</p>",
              "rawMarkdown": "Do they annotate Mass or Cancer? Are they same thing?"
            },
            {
              "id": 2070488,
              "postDate": "2022-12-20T04:06:05.810Z",
              "content": "<p>breast-level_annotations.csv: Each row corresponds to an image and provides the BI-RADS assessment of the breast depicted by the image along with some metadata of the image. The attributes in each row are:<br>\nstudy_id: The encoded study identifier.<br>\nseries_id: The encoded series identifier.<br>\nlaterality: Laterality of the breast depicted in the image. Either L or R.<br>\nview_position: Orientation with respect to the breast of the image. Standard views are CC and MLO.<br>\nheight: Height of the image.<br>\nwidth: Width of the image.<br>\nbreast_birads: BI-RADS assessment of the breast that the image depicts.<br>\nbreast_density: Density category of the breast that the image depicts.<br>\nsplit: indicating the split to which the image belongs, either training or test.<br>\nfinding_annotations.csv: Each row represents an annotation of a breast abnormality in an image. Metadata for each finding annotation includes image's metadata, namely image_id, study_id, series_id, laterality, view_positition, height, width, breast_birads, breast_density, and split, and annotation's metadata:<br>\nfinding_categories: List of finding categories attached to the marked region. For example, mass with skin retraction would be represented as [\"Mass\", \"Skin Retraction\"].<br>\nfinding_birads: BI-RADS assessment of the marked finding.<br>\nxmin: Left boundary of the box.<br>\nymin: Top boundary of the box.<br>\nxmax: Right boundary of the box.<br>\nymax: Bottom boundary of the box.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F0bc760b3e3a7e65249549e797645814a%2FScreenshot%202022-12-19%208.04.23%20PM.png?generation=1671509123741294&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "breast-level_annotations.csv: Each row corresponds to an image and provides the BI-RADS assessment of the breast depicted by the image along with some metadata of the image. The attributes in each row are:\nstudy_id: The encoded study identifier.\nseries_id: The encoded series identifier.\nlaterality: Laterality of the breast depicted in the image. Either L or R.\nview_position: Orientation with respect to the breast of the image. Standard views are CC and MLO.\nheight: Height of the image.\nwidth: Width of the image.\nbreast_birads: BI-RADS assessment of the breast that the image depicts.\nbreast_density: Density category of the breast that the image depicts.\nsplit: indicating the split to which the image belongs, either training or test.\nfinding_annotations.csv: Each row represents an annotation of a breast abnormality in an image. Metadata for each finding annotation includes image's metadata, namely image_id, study_id, series_id, laterality, view_positition, height, width, breast_birads, breast_density, and split, and annotation's metadata:\nfinding_categories: List of finding categories attached to the marked region. For example, mass with skin retraction would be represented as [\"Mass\", \"Skin Retraction\"].\nfinding_birads: BI-RADS assessment of the marked finding.\nxmin: Left boundary of the box.\nymin: Top boundary of the box.\nxmax: Right boundary of the box.\nymax: Bottom boundary of the box.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F0bc760b3e3a7e65249549e797645814a%2FScreenshot%202022-12-19%208.04.23%20PM.png?generation=1671509123741294&alt=media)"
            }
          ]
        },
        {
          "id": 2072442,
          "postDate": "2022-12-22T05:38:39.153Z",
          "content": "<p>there is one trick here</p>\n<ul>\n<li>the competition is about image level</li>\n<li>so your object detector yolo is allowed to make false positive (or even false negative) in the image if i can detect at least one true positive.</li>\n<li>this is why image level accuracy is better (much better) then roi accuracy.</li>\n</ul>\n<hr>\n<p>actually our competition is not image level it is patient laterality level. so it can make mistakes even at image level.</p>",
          "rawMarkdown": "there is one trick here\n- the competition is about image level\n- so your object detector yolo is allowed to make false positive (or even false negative) in the image if i can detect at least one true positive.\n- this is why image level accuracy is better (much better) then roi accuracy.\n\n---\n\nactually our competition is not image level it is patient laterality level. so it can make mistakes even at image level.",
          "replies": [
            {
              "id": 2072470,
              "postDate": "2022-12-22T06:03:26.563Z",
              "content": "<p>Yeah image level is a rather weird, because it's really unusable in any diagonistic setting.  As a pre-processor image level might be sort of useful?  I dunno</p>\n<p>But yes, FP are absolutely expected with ROI extractors and in fact encouraged, which is why you reduce the confidence thresholds.  </p>\n<p>btw, FP in general is <em>extremely</em> high in BC scanning, like 95%.  With FN unfortunately hovering around 12.5%, generally you just want to figure out whether or not to get a biospy </p>\n<p>once you have a bunch of FP + hopefully TP ROI, than you drill down via a more focused model</p>",
              "rawMarkdown": "Yeah image level is a rather weird, because it's really unusable in any diagonistic setting.  As a pre-processor image level might be sort of useful?  I dunno\n\nBut yes, FP are absolutely expected with ROI extractors and in fact encouraged, which is why you reduce the confidence thresholds.  \n\nbtw, FP in general is *extremely* high in BC scanning, like 95%.  With FN unfortunately hovering around 12.5%, generally you just want to figure out whether or not to get a biospy \n\nonce you have a bunch of FP + hopefully TP ROI, than you drill down via a more focused model\n\n"
            },
            {
              "id": 2072477,
              "postDate": "2022-12-22T06:12:43.830Z",
              "content": "<p>my suggestion is not to focus on processing. train a multiview classifier at the end (e.g. multi-instance) to make prediction at the patient laterality level.</p>",
              "rawMarkdown": "my suggestion is not to focus on processing. train a multiview classifier at the end (e.g. multi-instance) to make prediction at the patient laterality level."
            },
            {
              "id": 2072499,
              "postDate": "2022-12-22T06:45:58.387Z",
              "content": "<p>Hmmm not finding many examples of mv in the literature (i prefer to search only in 2022), but using a mv classifier is an interesting idea.  Do you have any cites that use ROI patch extractors, preferably yolo?</p>\n<p><a href=\"https://scholar.google.com/scholar?hl=en&amp;as_sdt=0%2C5&amp;as_ylo=2022&amp;q=%22multi-view+classifier%22+%22breast+cancer%22&amp;btnG=\" target=\"_blank\">https://scholar.google.com/scholar?hl=en&amp;as_sdt=0%2C5&amp;as_ylo=2022&amp;q=%22multi-view+classifier%22+%22breast+cancer%22&amp;btnG=</a></p>\n<p>I'm thinking more of a collection of patch classifier models, with issues being how to propagate back from the classifier to the beginning of yolo training</p>\n<p>so image -&gt; a) ROI patch extractor -&gt; b) patch classifier, and some mechanism to propagate back gradients to a)</p>\n<p>note that yolo isn't the only way to do ROI patch extraction.  wavelets are really interesting as well, also very intersting for patch classification</p>",
              "rawMarkdown": "Hmmm not finding many examples of mv in the literature (i prefer to search only in 2022), but using a mv classifier is an interesting idea.  Do you have any cites that use ROI patch extractors, preferably yolo?\n  \nhttps://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&as_ylo=2022&q=%22multi-view+classifier%22+%22breast+cancer%22&btnG=\n\nI'm thinking more of a collection of patch classifier models, with issues being how to propagate back from the classifier to the beginning of yolo training\n\nso image -> a) ROI patch extractor -> b) patch classifier, and some mechanism to propagate back gradients to a)\n\nnote that yolo isn't the only way to do ROI patch extraction.  wavelets are really interesting as well, also very intersting for patch classification"
            },
            {
              "id": 2072504,
              "postDate": "2022-12-22T06:49:55.260Z",
              "content": "<p>One of the things is I'd like to make a model that also works for ROI regions, tbh.  Scoring well in this comp is a primary goal, of course, but I'm also interested in making something generally useful.  It's very freeing when your'e not trying to win, just contribute useful things for folks via discussions and notebooks.  It's this balance that makes kaggle so very cool</p>",
              "rawMarkdown": "One of the things is I'd like to make a model that also works for ROI regions, tbh.  Scoring well in this comp is a primary goal, of course, but I'm also interested in making something generally useful.  It's very freeing when your'e not trying to win, just contribute useful things for folks via discussions and notebooks.  It's this balance that makes kaggle so very cool"
            }
          ]
        }
      ]
    },
    {
      "id": 2065437,
      "postDate": "2022-12-14T16:52:19.880Z",
      "content": "<p>that's something definetely worth trying, either with the mask or the bbox.</p>",
      "rawMarkdown": "that's something definetely worth trying, either with the mask or the bbox.",
      "votes": 1
    },
    {
      "id": 2066698,
      "postDate": "2022-12-16T00:59:27.303Z",
      "content": "<p>Awsaf, how realiable is the grad cam?  Wouldn't it be better to utilize your DDSM dataset you uploaded to kaggle?    </p>\n<p><a href=\"https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset</a></p>\n<p>eg:  <a href=\"http://www.eng.usf.edu/cvprg/Mammography/Database.html\" target=\"_blank\">http://www.eng.usf.edu/cvprg/Mammography/Database.html</a></p>\n<p>Also, what do you think about looking at the difference between right/left laterality?  If you look at train, it's fairly rare for cancer to be in both sides.</p>\n<p>eg:</p>\n<p><a href=\"http://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html\" target=\"_blank\">http://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html</a></p>",
      "rawMarkdown": "Awsaf, how realiable is the grad cam?  Wouldn't it be better to utilize your DDSM dataset you uploaded to kaggle?    \n\nhttps://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\n\neg:  http://www.eng.usf.edu/cvprg/Mammography/Database.html\n\nAlso, what do you think about looking at the difference between right/left laterality?  If you look at train, it's fairly rare for cancer to be in both sides.\n\neg:\n\nhttp://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html",
      "votes": 2,
      "replies": [
        {
          "id": 2066787,
          "postDate": "2022-12-16T04:13:18.280Z",
          "content": "<p>It would be great to try <strong>DDSM</strong> but there is a chance that competition dataset and <strong>DDSM</strong> may have some processing differences. It would be nice to try. </p>",
          "rawMarkdown": "It would be great to try **DDSM** but there is a chance that competition dataset and **DDSM** may have some processing differences. It would be nice to try. ",
          "votes": 1,
          "replies": [
            {
              "id": 2068827,
              "postDate": "2022-12-18T10:47:07.477Z",
              "content": "<p>Maybe we need to do a combination of different approaches, ddsm, grad cam, aother salient annotation sets.   But I suspect yolo will be the object detector.</p>",
              "rawMarkdown": "Maybe we need to do a combination of different approaches, ddsm, grad cam, aother salient annotation sets.   But I suspect yolo will be the object detector."
            }
          ]
        }
      ]
    },
    {
      "id": 2789456,
      "postDate": "2024-05-02T17:18:36.137Z",
      "content": "<p>I Recently gone through How to train Neural Network , It was an Awesome Experience  were i loaded Breast Cancer Dataset using scikit learn , It was an amazing Experience </p>",
      "rawMarkdown": "I Recently gone through How to train Neural Network , It was an Awesome Experience  were i loaded Breast Cancer Dataset using scikit learn , It was an amazing Experience "
    },
    {
      "id": 2076704,
      "postDate": "2022-12-26T18:09:20.127Z",
      "content": "<p>A Study on Class Activation Map Methods to Detect Masses in<br>\nMammography Images using Weakly Supervised Learning</p>\n<p><a href=\"http://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf\" target=\"_blank\">http://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf</a></p>",
      "rawMarkdown": "A Study on Class Activation Map Methods to Detect Masses in\nMammography Images using Weakly Supervised Learning\n\nhttp://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf"
    },
    {
      "id": 2065569,
      "postDate": "2022-12-14T20:23:10.750Z",
      "content": "<p>How reliable the prediction of a model with such highly imbalance dateset? Also, domain expert is needed to consider the right prediction. Otherwise, this method might be risky. </p>",
      "rawMarkdown": "How reliable the prediction of a model with such highly imbalance dateset? Also, domain expert is needed to consider the right prediction. Otherwise, this method might be risky. "
    }
  ],
  "comments": [
    {
      "id": 2067791,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-17T06:48:46.457000",
      "content": "<p><img src=\"https://i.ibb.co/f8YBSgd/mia-structure.png\" alt=\"https://i.ibb.co/f8YBSgd/mia-structure.png\"><br>\n<a href=\"https://github.com/nyukat/GMIC\" target=\"_blank\">https://github.com/nyukat/GMIC</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2065432,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-14T16:49:00.170000",
      "content": "<p>a better solution is:</p>\n<ol>\n<li>find open source dataset with lesion annotation (there are many of them, some are box annotations and some are mask annotations) i think you can get about 3k to 4k annotations.</li>\n<li>train a segmentation model using those external dataset</li>\n<li>apply the segmentation model on kaggle dataset to make output predicted probability as a new channel </li>\n<li>now you have input=gray+probability to train a new kaggle model.</li>\n</ol>\n<p>shortcut:</p>\n<ul>\n<li>use opensource segmentation model and avoid step.2 for a quick test.<br>\ne.g. <a href=\"https://github.com/nyukat/\" target=\"_blank\">https://github.com/nyukat/</a></li>\n</ul>\n<p>these models are trained on millions of mammography images.<br>\nthe approach described above is also based on their papers on the github</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2065440,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-14T16:53:57.733000",
          "content": "<p>just curious  …<br>\n<img src=\"https://i.ibb.co/vq5F86x/Selection-196.png\" alt=\"https://i.ibb.co/vq5F86x/Selection-196.png\"></p>\n<p>actually, GPTchat is good at answering general machine-learning problems. i have asked other questions regarding class imbalance, etc, and GPTchat gives good answers … maybe because he is just talking about his own personal experiences, how he is trained by openAI</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2067762,
              "author_name": "Haw Keat",
              "author_url": "",
              "post_date": "2022-12-17T06:00:08.557000",
              "content": "<p>Did GPTChat write the code for you? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2070287,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-19T19:57:03.957000",
      "content": "<p>I created a yolov5 POC here .. for the test I used the same images as the grad cam notebook listed above.</p>\n<p><a href=\"https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\" target=\"_blank\">https://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fd174d1dfa9098c927d7b8eb95f422c04%2FScreenshot%202022-12-19%2011.54.30%20AM.png?generation=1671479701605290&amp;alt=media\" alt=\"\"></p>\n<p>It's using the yolov5 models trained here - <a href=\"https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0\" target=\"_blank\">https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0</a></p>\n<p>I reduced the conf thresh to 0.1 .. other params can be tweaked.  Overlapping boundary boxes will probably want to be merged when pushing them through the next stage of the model (efnet, resnet, etc that detects benign/cancer).  Likely the yolo models need more training as well, but it's a start.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2070457,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-20T02:54:17.333000",
          "content": "<p>That's great =) Try pseudo labeling to generate labels for restrest of the positive images. Then re-train.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2070471,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-20T03:32:06.990000",
              "content": "<p><a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> </p>\n<p>I think the right idea is to go hard on </p>\n<p>Pham, H. H., Nguyen Trung, H., &amp; Nguyen, H. Q. (2022). VinDr-Mammo: A large-scale benchmark dataset for computer-aided detection and diagnosis in full-field digital mammography (version 1.0.0). PhysioNet. <a href=\"https://doi.org/10.13026/br2v-7517\" target=\"_blank\">https://doi.org/10.13026/br2v-7517</a>.</p>\n<p>It's FFDM and it's huge (5000 patients), and it has finding_annotations.csv to train yolo.</p>\n<p>But for folks competing, they might not be able to. I've reached out to the physionet folks to see if it would be permissible.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070483,
              "author_name": "Awsaf",
              "author_url": "",
              "post_date": "2022-12-20T04:00:03.897000",
              "content": "<p>Do they annotate Mass or Cancer? Are they same thing?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2070488,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-20T04:06:05.810000",
              "content": "<p>breast-level_annotations.csv: Each row corresponds to an image and provides the BI-RADS assessment of the breast depicted by the image along with some metadata of the image. The attributes in each row are:<br>\nstudy_id: The encoded study identifier.<br>\nseries_id: The encoded series identifier.<br>\nlaterality: Laterality of the breast depicted in the image. Either L or R.<br>\nview_position: Orientation with respect to the breast of the image. Standard views are CC and MLO.<br>\nheight: Height of the image.<br>\nwidth: Width of the image.<br>\nbreast_birads: BI-RADS assessment of the breast that the image depicts.<br>\nbreast_density: Density category of the breast that the image depicts.<br>\nsplit: indicating the split to which the image belongs, either training or test.<br>\nfinding_annotations.csv: Each row represents an annotation of a breast abnormality in an image. Metadata for each finding annotation includes image's metadata, namely image_id, study_id, series_id, laterality, view_positition, height, width, breast_birads, breast_density, and split, and annotation's metadata:<br>\nfinding_categories: List of finding categories attached to the marked region. For example, mass with skin retraction would be represented as [\"Mass\", \"Skin Retraction\"].<br>\nfinding_birads: BI-RADS assessment of the marked finding.<br>\nxmin: Left boundary of the box.<br>\nymin: Top boundary of the box.<br>\nxmax: Right boundary of the box.<br>\nymax: Bottom boundary of the box.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2F0bc760b3e3a7e65249549e797645814a%2FScreenshot%202022-12-19%208.04.23%20PM.png?generation=1671509123741294&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2072442,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-22T05:38:39.153000",
          "content": "<p>there is one trick here</p>\n<ul>\n<li>the competition is about image level</li>\n<li>so your object detector yolo is allowed to make false positive (or even false negative) in the image if i can detect at least one true positive.</li>\n<li>this is why image level accuracy is better (much better) then roi accuracy.</li>\n</ul>\n<hr>\n<p>actually our competition is not image level it is patient laterality level. so it can make mistakes even at image level.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2072470,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-22T06:03:26.563000",
              "content": "<p>Yeah image level is a rather weird, because it's really unusable in any diagonistic setting.  As a pre-processor image level might be sort of useful?  I dunno</p>\n<p>But yes, FP are absolutely expected with ROI extractors and in fact encouraged, which is why you reduce the confidence thresholds.  </p>\n<p>btw, FP in general is <em>extremely</em> high in BC scanning, like 95%.  With FN unfortunately hovering around 12.5%, generally you just want to figure out whether or not to get a biospy </p>\n<p>once you have a bunch of FP + hopefully TP ROI, than you drill down via a more focused model</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072477,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-22T06:12:43.830000",
              "content": "<p>my suggestion is not to focus on processing. train a multiview classifier at the end (e.g. multi-instance) to make prediction at the patient laterality level.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072499,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-22T06:45:58.387000",
              "content": "<p>Hmmm not finding many examples of mv in the literature (i prefer to search only in 2022), but using a mv classifier is an interesting idea.  Do you have any cites that use ROI patch extractors, preferably yolo?</p>\n<p><a href=\"https://scholar.google.com/scholar?hl=en&amp;as_sdt=0%2C5&amp;as_ylo=2022&amp;q=%22multi-view+classifier%22+%22breast+cancer%22&amp;btnG=\" target=\"_blank\">https://scholar.google.com/scholar?hl=en&amp;as_sdt=0%2C5&amp;as_ylo=2022&amp;q=%22multi-view+classifier%22+%22breast+cancer%22&amp;btnG=</a></p>\n<p>I'm thinking more of a collection of patch classifier models, with issues being how to propagate back from the classifier to the beginning of yolo training</p>\n<p>so image -&gt; a) ROI patch extractor -&gt; b) patch classifier, and some mechanism to propagate back gradients to a)</p>\n<p>note that yolo isn't the only way to do ROI patch extraction.  wavelets are really interesting as well, also very intersting for patch classification</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2072504,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-22T06:49:55.260000",
              "content": "<p>One of the things is I'd like to make a model that also works for ROI regions, tbh.  Scoring well in this comp is a primary goal, of course, but I'm also interested in making something generally useful.  It's very freeing when your'e not trying to win, just contribute useful things for folks via discussions and notebooks.  It's this balance that makes kaggle so very cool</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2065437,
      "author_name": "Eleftherios Fanioudakis",
      "author_url": "",
      "post_date": "2022-12-14T16:52:19.880000",
      "content": "<p>that's something definetely worth trying, either with the mask or the bbox.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2066698,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-16T00:59:27.303000",
      "content": "<p>Awsaf, how realiable is the grad cam?  Wouldn't it be better to utilize your DDSM dataset you uploaded to kaggle?    </p>\n<p><a href=\"https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset</a></p>\n<p>eg:  <a href=\"http://www.eng.usf.edu/cvprg/Mammography/Database.html\" target=\"_blank\">http://www.eng.usf.edu/cvprg/Mammography/Database.html</a></p>\n<p>Also, what do you think about looking at the difference between right/left laterality?  If you look at train, it's fairly rare for cancer to be in both sides.</p>\n<p>eg:</p>\n<p><a href=\"http://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html\" target=\"_blank\">http://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2066787,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-16T04:13:18.280000",
          "content": "<p>It would be great to try <strong>DDSM</strong> but there is a chance that competition dataset and <strong>DDSM</strong> may have some processing differences. It would be nice to try. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2068827,
              "author_name": "@kaggleqrdl",
              "author_url": "",
              "post_date": "2022-12-18T10:47:07.477000",
              "content": "<p>Maybe we need to do a combination of different approaches, ddsm, grad cam, aother salient annotation sets.   But I suspect yolo will be the object detector.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2789456,
      "author_name": "Jayesh Kothavale",
      "author_url": "",
      "post_date": "2024-05-02T17:18:36.137000",
      "content": "<p>I Recently gone through How to train Neural Network , It was an Awesome Experience  were i loaded Breast Cancer Dataset using scikit learn , It was an amazing Experience </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2076704,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-26T18:09:20.127000",
      "content": "<p>A Study on Class Activation Map Methods to Detect Masses in<br>\nMammography Images using Weakly Supervised Learning</p>\n<p><a href=\"http://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf\" target=\"_blank\">http://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2065569,
      "author_name": "Simon Alerdic",
      "author_url": "",
      "post_date": "2022-12-14T20:23:10.750000",
      "content": "<p>How reliable the prediction of a model with such highly imbalance dateset? Also, domain expert is needed to consider the right prediction. Otherwise, this method might be risky. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2065413": "I am wondering if we can convert this **cancer detection** problem from `classification` to `object detection`. There are couple of good reasons for considering object detection, such as cancer pixels (object) contains very small portion of the image. If we consider full dicom (w/o roi) then it becomes even worse. The model has to overcome a lot of distractions due to such imbalance. In object detection, model can optimize more easily using **bounding box**. Then there are some models like **yolov5**, **yolov7** and so on..\n\nBut sadly we don't have any **bounding box**. But we do have access to **class activation map (Grad-CAM)** which shows parts of an input image that most impact the classification score. What if we convert the  **grad-cam** to **bounding box (bbox)** ? I published following notebook to carry out this experiment with **WandB**.\n\n## Notebook:\n* [RSNA-BCD: GradCAM to BBox](https://www.kaggle.com/code/awsaf49/rsna-bcd-gradcam-to-bbox)\n* RSNA Breast Cancer Detection [YOLOv5] (Coming Soon...)\n\n## Steps:\n1. Take **cancer** images only.\n2. Apply filter to **cancer** images to take only images with **good confidence**. The threshold  for `confidence` can be obtained by maximizing `pF1` score like we're doing in classification.\n4. Convert **gradcam** to **mask** using `thresholding` then obtain **bbox** from **mask** using `largest contour`, so in a nutshell, image -> gradcam -> mask -> bbox\n3. In training, use only the **filtered cancer** images obtained from the above steps.\n4. Apply pseudo labeling on other **cancer** images to obtain their labels.\n\n## WandB\n<span style=\"color: #000508; font-family: Segoe UI; font-size: 1.5em; font-weight: 300;\"><a href=\"https://wandb.ai/awsaf49/rsna-bcd-gradcam2bbox\">View the Complete Dashboard Here ⮕</a></span>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb96243b27b03d4635f6a9242b2c29372%2Fgradcam2bbox-wandb.PNG?generation=1671034084366419&alt=media)\n\n## Samples\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F31e9e3bd9915e6fc120c79f212d366e2%2Fgradcam2bbox.jpg?generation=1671033517214365&alt=media\" height=\"900\">\n\n## Key-Points\n* GradCAM doesn't always point to the right location even when the prediction is correct. So, we need to use this method very carefully. Usually, high confidence means grad-cam is more likely to point to the correct location.\n* Better classification models definitely generate better bounding box, hence need accurate models.\n* We can ensemble different mdoels' gradcams by using method like **Weighted Boxes Fusion (WBF)** by considering model's prediction as confidence score of **bounding box**.",
    "2067791": "![https://i.ibb.co/f8YBSgd/mia-structure.png](https://i.ibb.co/f8YBSgd/mia-structure.png)\nhttps://github.com/nyukat/GMIC",
    "2065432": "a better solution is:\n1.  find open source dataset with lesion annotation (there are many of them, some are box annotations and some are mask annotations) i think you can get about 3k to 4k annotations.\n2. train a segmentation model using those external dataset\n3. apply the segmentation model on kaggle dataset to make output predicted probability as a new channel \n4. now you have input=gray+probability to train a new kaggle model.\n\nshortcut:\n- use opensource segmentation model and avoid step.2 for a quick test.\ne.g. https://github.com/nyukat/\n\nthese models are trained on millions of mammography images.\nthe approach described above is also based on their papers on the github",
    "2070287": "I created a yolov5 POC here .. for the test I used the same images as the grad cam notebook listed above.\n\nhttps://www.kaggle.com/code/kaggleqrdl/yolov5-roi-patch-poc\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9052057%2Fd174d1dfa9098c927d7b8eb95f422c04%2FScreenshot%202022-12-19%2011.54.30%20AM.png?generation=1671479701605290&alt=media)\n\nIt's using the yolov5 models trained here - https://github.com/vbookshelf/Mammogram-Mass-Analyzer/blob/main/mammogram-mass-analyzer-v0.0\n\nI reduced the conf thresh to 0.1 .. other params can be tweaked.  Overlapping boundary boxes will probably want to be merged when pushing them through the next stage of the model (efnet, resnet, etc that detects benign/cancer).  Likely the yolo models need more training as well, but it's a start.",
    "2065437": "that's something definetely worth trying, either with the mask or the bbox.",
    "2066698": "Awsaf, how realiable is the grad cam?  Wouldn't it be better to utilize your DDSM dataset you uploaded to kaggle?    \n\nhttps://www.kaggle.com/datasets/awsaf49/cbis-ddsm-breast-cancer-image-dataset\n\neg:  http://www.eng.usf.edu/cvprg/Mammography/Database.html\n\nAlso, what do you think about looking at the difference between right/left laterality?  If you look at train, it's fairly rare for cancer to be in both sides.\n\neg:\n\nhttp://www.eng.usf.edu/cvprg/Mammography/DDSM/thumbnails/cancers/cancer_01/case0001/C-0001-1.html",
    "2789456": "I Recently gone through How to train Neural Network , It was an Awesome Experience  were i loaded Breast Cancer Dataset using scikit learn , It was an amazing Experience ",
    "2076704": "A Study on Class Activation Map Methods to Detect Masses in\nMammography Images using Weakly Supervised Learning\n\nhttp://sites.labic.icmc.usp.br/eniac2022/pdf/38.pdf",
    "2065569": "How reliable the prediction of a model with such highly imbalance dateset? Also, domain expert is needed to consider the right prediction. Otherwise, this method might be risky. "
  }
}