{
  "id": 371077,
  "title": "Caution: Enriched dataset",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/371077",
  "author_name": "@kaggleqrdl",
  "post_date": "2022-12-07T20:45:01.566000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> dropped this nugget on one of the threads</p>\n<blockquote>\n  <p>However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107</a></p>\n<p>Indeed, I noticed a statistical curiosity.   Could be coincidence, but if not certainly would point to an enriched dataset.</p>\n<pre><code>gb_lat = origtrain.drop_duplicates([, ]).groupby([, , ]).count()[]\ndisplay(gb_lat)\n</code></pre>\n<pre><code>laterality  site_id  cancer\nL                            \n                               \n                             \n                               \nR                            \n                               \n                             \n                               \nName: age, dtype: int64\n</code></pre>\n<p>My reading of this is that site 2 has the same amount of Left and Right laterality incidences of cancer.  It could be total coincidence, of course, but it could also be an example of how the data managers wanted to make sure there were sufficient examples of both laterality with cancer.  </p>\n<p>Literature is a bit conflicted on this, but some reports say that left breasts have a 5-10% greater chance of cancer.  The question of course is how enriched the hidden set is.</p>",
  "messages": [
    {
      "id": 2058352,
      "postDate": "2022-12-07T20:45:01.567Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> dropped this nugget on one of the threads</p>\n<blockquote>\n  <p>However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107</a></p>\n<p>Indeed, I noticed a statistical curiosity.   Could be coincidence, but if not certainly would point to an enriched dataset.</p>\n<pre><code>gb_lat = origtrain.drop_duplicates([, ]).groupby([, , ]).count()[]\ndisplay(gb_lat)\n</code></pre>\n<pre><code>laterality  site_id  cancer\nL                            \n                               \n                             \n                               \nR                            \n                               \n                             \n                               \nName: age, dtype: int64\n</code></pre>\n<p>My reading of this is that site 2 has the same amount of Left and Right laterality incidences of cancer.  It could be total coincidence, of course, but it could also be an example of how the data managers wanted to make sure there were sufficient examples of both laterality with cancer.  </p>\n<p>Literature is a bit conflicted on this, but some reports say that left breasts have a 5-10% greater chance of cancer.  The question of course is how enriched the hidden set is.</p>",
      "rawMarkdown": "@sohier dropped this nugget on one of the threads\n\n> However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\n\nIndeed, I noticed a statistical curiosity.   Could be coincidence, but if not certainly would point to an enriched dataset.\n\n```python\ngb_lat = origtrain.drop_duplicates(['patient_id', 'laterality']).groupby([\"laterality\", \"site_id\", 'cancer']).count()['age']\ndisplay(gb_lat)\n```\n\n```python\nlaterality  site_id  cancer\nL           1        0         5681\n                     1          129\n            2        0         5976\n                     1          119\nR           1        0         5685\n                     1          125\n            2        0         5976\n                     1          119\nName: age, dtype: int64\n```\n\nMy reading of this is that site 2 has the same amount of Left and Right laterality incidences of cancer.  It could be total coincidence, of course, but it could also be an example of how the data managers wanted to make sure there were sufficient examples of both laterality with cancer.  \n\nLiterature is a bit conflicted on this, but some reports say that left breasts have a 5-10% greater chance of cancer.  The question of course is how enriched the hidden set is.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2058352": "@sohier dropped this nugget on one of the threads\n\n> However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107\n\nIndeed, I noticed a statistical curiosity.   Could be coincidence, but if not certainly would point to an enriched dataset.\n\n```python\ngb_lat = origtrain.drop_duplicates(['patient_id', 'laterality']).groupby([\"laterality\", \"site_id\", 'cancer']).count()['age']\ndisplay(gb_lat)\n```\n\n```python\nlaterality  site_id  cancer\nL           1        0         5681\n                     1          129\n            2        0         5976\n                     1          119\nR           1        0         5685\n                     1          125\n            2        0         5976\n                     1          119\nName: age, dtype: int64\n```\n\nMy reading of this is that site 2 has the same amount of Left and Right laterality incidences of cancer.  It could be total coincidence, of course, but it could also be an example of how the data managers wanted to make sure there were sufficient examples of both laterality with cancer.  \n\nLiterature is a bit conflicted on this, but some reports say that left breasts have a 5-10% greater chance of cancer.  The question of course is how enriched the hidden set is."
  }
}