{
  "id": 369660,
  "title": "Mirai - Mammography-based model for breast cancer risk",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369660",
  "author_name": "Darien Schettler",
  "post_date": "2022-11-30T23:03:36.932000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I wanted to share this resource as it appears to be one of the most SOTA recent machine learning approaches towards predicting Breast Cancer from Mammograms. However, I don't believe it's directly applicable/transferable (although you can spin up a server) due to the more stringent requirements and differences in goal/data (you need 4 views, etc.). </p>\n<p>That being said, there is enough overlap with our goal that it's not unreasonable to read through the paper and codebase and take inspiration.</p>\n<p>Here are the relevant links:</p>\n<ul>\n<li><a href=\"https://github.com/yala/Mirai\" target=\"_blank\"><strong>GITHUB</strong>: https://github.com/yala/Mirai</a></li>\n<li><a href=\"https://www.science.org/doi/10.1126/scitranslmed.aba4373\" target=\"_blank\"><strong>PAPER</strong>: Towards Robust Mammography-Based Models for Breast Cancer Risk</a></li>\n</ul>\n<p><br></p>\n<p>Here is the description from the Github ReadMe:</p>\n<blockquote>\n  <p>This repository was used to develop Mirai, the risk model described in: Towards Robust Mammography-Based Models for Breast Cancer Risk. Mirai was designed to predict risk at multiple time points, leverage potentially missing risk-factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the US and was tested on held-out test sets from MGH, Karolinska in Sweden and Chang Gung Memorial Hospital in Taiwan, obtaining C-indices of 0.76 (0.74, 0.80), 0.81 (0.79, 0.82), 0.79 (0.79, 0.83), respectively. Mirai obtained significantly higher five-year ROC AUCs than the Tyrer-Cuzick model (p&lt;0.001) and prior deep learning models, Hybrid DL (p&lt;0.001) and ImageOnly DL (p&lt;0.001), trained on the same MGH dataset. In our paper, we also demonstrate that Mirai was more significantly accurate in identifying high risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4, 48.5) of patients who would develop cancer within five-years were identified as high risk, compared to 36.1% (29.1, 42.9) by Hybrid DL (p=0.02) and 22.9% (15.9, 29.6) by Tyrer-Cuzick lifetime risk (p&lt;0.001).</p>\n</blockquote>\n<p><br></p>\n<p>Here is the abstract from the Paper:</p>\n<blockquote>\n  <p>Improved breast cancer risk models enable targeted screening strategies that achieve earlier detection and less screening harm than existing guidelines. To bring deep learning risk models to clinical practice, we need to further refine their accuracy, validate them across diverse populations, and demonstrate their potential to improve clinical workflows. We developed Mirai, a mammography-based deep learning model designed to predict risk at multiple timepoints, leverage potentially missing risk factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the United States and tested on held-out test sets from MGH, Karolinska University Hospital in Sweden, and Chang Gung Memorial Hospital (CGMH) in Taiwan, obtaining C-indices of 0.76 (95% confidence interval, 0.74 to 0.80), 0.81 (0.79 to 0.82), and 0.79 (0.79 to 0.83), respectively. Mirai obtained significantly higher 5-year ROC AUCs than the Tyrer-Cuzick model (P &lt; 0.001) and prior deep learning models Hybrid DL (P &lt; 0.001) and Image-Only DL (P &lt; 0.001), trained on the same dataset. Mirai more accurately identified high-risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4 to 48.5) of patients who would develop cancer within 5 years were identified as high risk, compared with 36.1% (29.1 to 42.9) by Hybrid DL (P = 0.02) and 22.9% (15.9 to 29.6) by the Tyrer-Cuzick model (P &lt; 0.001).</p>\n</blockquote>\n<hr>\n<p>I hope this helps.</p>\n<p>ps: It's a bit frustrating we can't access the same rich dataset that these researchers did. While it is the largest dataset of this type… it appears to be unavailable to the public</p>",
  "messages": [
    {
      "id": 2050664,
      "postDate": "2022-11-30T23:03:36.933Z",
      "content": "<p>I wanted to share this resource as it appears to be one of the most SOTA recent machine learning approaches towards predicting Breast Cancer from Mammograms. However, I don't believe it's directly applicable/transferable (although you can spin up a server) due to the more stringent requirements and differences in goal/data (you need 4 views, etc.). </p>\n<p>That being said, there is enough overlap with our goal that it's not unreasonable to read through the paper and codebase and take inspiration.</p>\n<p>Here are the relevant links:</p>\n<ul>\n<li><a href=\"https://github.com/yala/Mirai\" target=\"_blank\"><strong>GITHUB</strong>: https://github.com/yala/Mirai</a></li>\n<li><a href=\"https://www.science.org/doi/10.1126/scitranslmed.aba4373\" target=\"_blank\"><strong>PAPER</strong>: Towards Robust Mammography-Based Models for Breast Cancer Risk</a></li>\n</ul>\n<p><br></p>\n<p>Here is the description from the Github ReadMe:</p>\n<blockquote>\n  <p>This repository was used to develop Mirai, the risk model described in: Towards Robust Mammography-Based Models for Breast Cancer Risk. Mirai was designed to predict risk at multiple time points, leverage potentially missing risk-factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the US and was tested on held-out test sets from MGH, Karolinska in Sweden and Chang Gung Memorial Hospital in Taiwan, obtaining C-indices of 0.76 (0.74, 0.80), 0.81 (0.79, 0.82), 0.79 (0.79, 0.83), respectively. Mirai obtained significantly higher five-year ROC AUCs than the Tyrer-Cuzick model (p&lt;0.001) and prior deep learning models, Hybrid DL (p&lt;0.001) and ImageOnly DL (p&lt;0.001), trained on the same MGH dataset. In our paper, we also demonstrate that Mirai was more significantly accurate in identifying high risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4, 48.5) of patients who would develop cancer within five-years were identified as high risk, compared to 36.1% (29.1, 42.9) by Hybrid DL (p=0.02) and 22.9% (15.9, 29.6) by Tyrer-Cuzick lifetime risk (p&lt;0.001).</p>\n</blockquote>\n<p><br></p>\n<p>Here is the abstract from the Paper:</p>\n<blockquote>\n  <p>Improved breast cancer risk models enable targeted screening strategies that achieve earlier detection and less screening harm than existing guidelines. To bring deep learning risk models to clinical practice, we need to further refine their accuracy, validate them across diverse populations, and demonstrate their potential to improve clinical workflows. We developed Mirai, a mammography-based deep learning model designed to predict risk at multiple timepoints, leverage potentially missing risk factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the United States and tested on held-out test sets from MGH, Karolinska University Hospital in Sweden, and Chang Gung Memorial Hospital (CGMH) in Taiwan, obtaining C-indices of 0.76 (95% confidence interval, 0.74 to 0.80), 0.81 (0.79 to 0.82), and 0.79 (0.79 to 0.83), respectively. Mirai obtained significantly higher 5-year ROC AUCs than the Tyrer-Cuzick model (P &lt; 0.001) and prior deep learning models Hybrid DL (P &lt; 0.001) and Image-Only DL (P &lt; 0.001), trained on the same dataset. Mirai more accurately identified high-risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4 to 48.5) of patients who would develop cancer within 5 years were identified as high risk, compared with 36.1% (29.1 to 42.9) by Hybrid DL (P = 0.02) and 22.9% (15.9 to 29.6) by the Tyrer-Cuzick model (P &lt; 0.001).</p>\n</blockquote>\n<hr>\n<p>I hope this helps.</p>\n<p>ps: It's a bit frustrating we can't access the same rich dataset that these researchers did. While it is the largest dataset of this type… it appears to be unavailable to the public</p>",
      "rawMarkdown": "I wanted to share this resource as it appears to be one of the most SOTA recent machine learning approaches towards predicting Breast Cancer from Mammograms. However, I don't believe it's directly applicable/transferable (although you can spin up a server) due to the more stringent requirements and differences in goal/data (you need 4 views, etc.). \n\nThat being said, there is enough overlap with our goal that it's not unreasonable to read through the paper and codebase and take inspiration.\n\nHere are the relevant links:\n* [**GITHUB**: https://github.com/yala/Mirai](https://github.com/yala/Mirai)\n* [**PAPER**: Towards Robust Mammography-Based Models for Breast Cancer Risk](https://www.science.org/doi/10.1126/scitranslmed.aba4373)\n\n<br>\n\nHere is the description from the Github ReadMe:\n\n> This repository was used to develop Mirai, the risk model described in: Towards Robust Mammography-Based Models for Breast Cancer Risk. Mirai was designed to predict risk at multiple time points, leverage potentially missing risk-factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the US and was tested on held-out test sets from MGH, Karolinska in Sweden and Chang Gung Memorial Hospital in Taiwan, obtaining C-indices of 0.76 (0.74, 0.80), 0.81 (0.79, 0.82), 0.79 (0.79, 0.83), respectively. Mirai obtained significantly higher five-year ROC AUCs than the Tyrer-Cuzick model (p<0.001) and prior deep learning models, Hybrid DL (p<0.001) and ImageOnly DL (p<0.001), trained on the same MGH dataset. In our paper, we also demonstrate that Mirai was more significantly accurate in identifying high risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4, 48.5) of patients who would develop cancer within five-years were identified as high risk, compared to 36.1% (29.1, 42.9) by Hybrid DL (p=0.02) and 22.9% (15.9, 29.6) by Tyrer-Cuzick lifetime risk (p<0.001).\n\n<br>\n\nHere is the abstract from the Paper:\n\n> Improved breast cancer risk models enable targeted screening strategies that achieve earlier detection and less screening harm than existing guidelines. To bring deep learning risk models to clinical practice, we need to further refine their accuracy, validate them across diverse populations, and demonstrate their potential to improve clinical workflows. We developed Mirai, a mammography-based deep learning model designed to predict risk at multiple timepoints, leverage potentially missing risk factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the United States and tested on held-out test sets from MGH, Karolinska University Hospital in Sweden, and Chang Gung Memorial Hospital (CGMH) in Taiwan, obtaining C-indices of 0.76 (95% confidence interval, 0.74 to 0.80), 0.81 (0.79 to 0.82), and 0.79 (0.79 to 0.83), respectively. Mirai obtained significantly higher 5-year ROC AUCs than the Tyrer-Cuzick model (P < 0.001) and prior deep learning models Hybrid DL (P < 0.001) and Image-Only DL (P < 0.001), trained on the same dataset. Mirai more accurately identified high-risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4 to 48.5) of patients who would develop cancer within 5 years were identified as high risk, compared with 36.1% (29.1 to 42.9) by Hybrid DL (P = 0.02) and 22.9% (15.9 to 29.6) by the Tyrer-Cuzick model (P < 0.001).\n\n---\n\nI hope this helps.\n\nps: It's a bit frustrating we can't access the same rich dataset that these researchers did. While it is the largest dataset of this type... it appears to be unavailable to the public",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2050664": "I wanted to share this resource as it appears to be one of the most SOTA recent machine learning approaches towards predicting Breast Cancer from Mammograms. However, I don't believe it's directly applicable/transferable (although you can spin up a server) due to the more stringent requirements and differences in goal/data (you need 4 views, etc.). \n\nThat being said, there is enough overlap with our goal that it's not unreasonable to read through the paper and codebase and take inspiration.\n\nHere are the relevant links:\n* [**GITHUB**: https://github.com/yala/Mirai](https://github.com/yala/Mirai)\n* [**PAPER**: Towards Robust Mammography-Based Models for Breast Cancer Risk](https://www.science.org/doi/10.1126/scitranslmed.aba4373)\n\n<br>\n\nHere is the description from the Github ReadMe:\n\n> This repository was used to develop Mirai, the risk model described in: Towards Robust Mammography-Based Models for Breast Cancer Risk. Mirai was designed to predict risk at multiple time points, leverage potentially missing risk-factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the US and was tested on held-out test sets from MGH, Karolinska in Sweden and Chang Gung Memorial Hospital in Taiwan, obtaining C-indices of 0.76 (0.74, 0.80), 0.81 (0.79, 0.82), 0.79 (0.79, 0.83), respectively. Mirai obtained significantly higher five-year ROC AUCs than the Tyrer-Cuzick model (p<0.001) and prior deep learning models, Hybrid DL (p<0.001) and ImageOnly DL (p<0.001), trained on the same MGH dataset. In our paper, we also demonstrate that Mirai was more significantly accurate in identifying high risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4, 48.5) of patients who would develop cancer within five-years were identified as high risk, compared to 36.1% (29.1, 42.9) by Hybrid DL (p=0.02) and 22.9% (15.9, 29.6) by Tyrer-Cuzick lifetime risk (p<0.001).\n\n<br>\n\nHere is the abstract from the Paper:\n\n> Improved breast cancer risk models enable targeted screening strategies that achieve earlier detection and less screening harm than existing guidelines. To bring deep learning risk models to clinical practice, we need to further refine their accuracy, validate them across diverse populations, and demonstrate their potential to improve clinical workflows. We developed Mirai, a mammography-based deep learning model designed to predict risk at multiple timepoints, leverage potentially missing risk factor information, and produce predictions that are consistent across mammography machines. Mirai was trained on a large dataset from Massachusetts General Hospital (MGH) in the United States and tested on held-out test sets from MGH, Karolinska University Hospital in Sweden, and Chang Gung Memorial Hospital (CGMH) in Taiwan, obtaining C-indices of 0.76 (95% confidence interval, 0.74 to 0.80), 0.81 (0.79 to 0.82), and 0.79 (0.79 to 0.83), respectively. Mirai obtained significantly higher 5-year ROC AUCs than the Tyrer-Cuzick model (P < 0.001) and prior deep learning models Hybrid DL (P < 0.001) and Image-Only DL (P < 0.001), trained on the same dataset. Mirai more accurately identified high-risk patients than prior methods across all datasets. On the MGH test set, 41.5% (34.4 to 48.5) of patients who would develop cancer within 5 years were identified as high risk, compared with 36.1% (29.1 to 42.9) by Hybrid DL (P = 0.02) and 22.9% (15.9 to 29.6) by the Tyrer-Cuzick model (P < 0.001).\n\n---\n\nI hope this helps.\n\nps: It's a bit frustrating we can't access the same rich dataset that these researchers did. While it is the largest dataset of this type... it appears to be unavailable to the public"
  }
}