{
  "id": 370341,
  "title": "Some LB probing results to share",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370341",
  "author_name": "tomoo inubushi",
  "post_date": "2022-12-04T01:16:20.284000",
  "votes": 67,
  "comment_count": 33,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about test dataset.<br>\nI demonstrated that following assumptions are all TRUE.</p>\n<ul>\n<li>There are no new site ID in test dataset.</li>\n<li>Patient IDs in train and test sets do not overlap</li>\n<li>Image IDs in train and test sets do not overlap</li>\n<li>There are no new laterality values in test dataset.</li>\n<li>There are new machine IDs in test dataset. (This is already raised by <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> in <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369362\" target=\"_blank\">here</a>)</li>\n<li>There are no new view values in test dataset.</li>\n<li>No. of images/patient are all &gt;= 4</li>\n<li>Site ID is always the same for each patient.</li>\n<li>Age is always the same for each patient.</li>\n<li>There are no overlap of machine IDs between two sites in test dataset.</li>\n<li>Some patients underwent mammography with multiple machines.</li>\n</ul>\n<p>[updated at 2022-12-04]</p>\n<ul>\n<li>No. of images in site ID 1 &gt; No. of images in site ID 2. (by <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>)</li>\n</ul>\n<p>[updated at 2022-12-05]</p>\n<ul>\n<li>All patients have CC and MLO images for both sides</li>\n<li>More than 40% of images are from machine ID 49 (43% for train dataset) (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li>Mean age of patients is between 56-61 (58.6 for train set), and patients in site 1 are younger than those in site 2. (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li>1-2% of patients use implants (1.4% for train set).(by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li><strong>Positive rate is between 0.0204-0.0256</strong>  (0.0206 for train dataset). (shown by <a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> and others)</li>\n</ul>\n<p>[updated at 2022-12-15]</p>\n<ul>\n<li>Age column contains nan, while others do not</li>\n</ul>\n<p>I am happy if anyone correct me if I am wrong.<br>\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.</p>",
  "messages": [
    {
      "id": 2054249,
      "postDate": "2022-12-04T01:16:20.283Z",
      "content": "<p>In <a href=\"https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook\" target=\"_blank\">this notebook</a>, I tested some basic assumptions about test dataset.<br>\nI demonstrated that following assumptions are all TRUE.</p>\n<ul>\n<li>There are no new site ID in test dataset.</li>\n<li>Patient IDs in train and test sets do not overlap</li>\n<li>Image IDs in train and test sets do not overlap</li>\n<li>There are no new laterality values in test dataset.</li>\n<li>There are new machine IDs in test dataset. (This is already raised by <a href=\"https://www.kaggle.com/abebe9849\" target=\"_blank\">@abebe9849</a> in <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369362\" target=\"_blank\">here</a>)</li>\n<li>There are no new view values in test dataset.</li>\n<li>No. of images/patient are all &gt;= 4</li>\n<li>Site ID is always the same for each patient.</li>\n<li>Age is always the same for each patient.</li>\n<li>There are no overlap of machine IDs between two sites in test dataset.</li>\n<li>Some patients underwent mammography with multiple machines.</li>\n</ul>\n<p>[updated at 2022-12-04]</p>\n<ul>\n<li>No. of images in site ID 1 &gt; No. of images in site ID 2. (by <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a>)</li>\n</ul>\n<p>[updated at 2022-12-05]</p>\n<ul>\n<li>All patients have CC and MLO images for both sides</li>\n<li>More than 40% of images are from machine ID 49 (43% for train dataset) (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li>Mean age of patients is between 56-61 (58.6 for train set), and patients in site 1 are younger than those in site 2. (by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li>1-2% of patients use implants (1.4% for train set).(by <a href=\"https://www.kaggle.com/kaggleqrdl\" target=\"_blank\">@kaggleqrdl</a>)</li>\n<li><strong>Positive rate is between 0.0204-0.0256</strong>  (0.0206 for train dataset). (shown by <a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> and others)</li>\n</ul>\n<p>[updated at 2022-12-15]</p>\n<ul>\n<li>Age column contains nan, while others do not</li>\n</ul>\n<p>I am happy if anyone correct me if I am wrong.<br>\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.</p>",
      "rawMarkdown": "In [this notebook](https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook), I tested some basic assumptions about test dataset.\nI demonstrated that following assumptions are all TRUE.\n\n* There are no new site ID in test dataset.\n* Patient IDs in train and test sets do not overlap\n* Image IDs in train and test sets do not overlap\n* There are no new laterality values in test dataset.\n* There are new machine IDs in test dataset. (This is already raised by @abebe9849 in [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369362))\n* There are no new view values in test dataset.\n* No. of images/patient are all >= 4\n* Site ID is always the same for each patient.\n* Age is always the same for each patient.\n* There are no overlap of machine IDs between two sites in test dataset.\n* Some patients underwent mammography with multiple machines.\n\n[updated at 2022-12-04]\n* No. of images in site ID 1 > No. of images in site ID 2. (by @yujiariyasu)\n\n[updated at 2022-12-05]\n* All patients have CC and MLO images for both sides\n* More than 40% of images are from machine ID 49 (43% for train dataset) (by @kaggleqrdl)\n* Mean age of patients is between 56-61 (58.6 for train set), and patients in site 1 are younger than those in site 2. (by @kaggleqrdl)\n* 1-2% of patients use implants (1.4% for train set).(by @kaggleqrdl)\n* **Positive rate is between 0.0204-0.0256**  (0.0206 for train dataset). (shown by @zzy990106 and others)\n\n[updated at 2022-12-15]\n* Age column contains nan, while others do not\n\nI am happy if anyone correct me if I am wrong.\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.",
      "votes": 67
    },
    {
      "id": 2148868,
      "postDate": "2023-02-17T18:14:19.580Z",
      "content": "<blockquote>\n  <p>All patients have CC and MLO images for both sides</p>\n</blockquote>\n<p>This is surprising, as there seemed to be examinations in the <strong>train set</strong> where at least one of the views were missing (e.g., 5 <code>DICOM</code> files, 3 L-CC, 1 L-MLO and 1 R-MLO, but no R-CC).</p>",
      "rawMarkdown": ">All patients have CC and MLO images for both sides\n\nThis is surprising, as there seemed to be examinations in the **train set** where at least one of the views were missing (e.g., 5 `DICOM` files, 3 L-CC, 1 L-MLO and 1 R-MLO, but no R-CC).",
      "votes": 1
    },
    {
      "id": 2101738,
      "postDate": "2023-01-16T06:26:37.303Z",
      "content": "<p>Thank you. Is this true: '<em>There are new machine IDs in test dataset'</em> ?</p>",
      "rawMarkdown": "Thank you. Is this true: '*There are new machine IDs in test dataset'* ?",
      "votes": 1,
      "replies": [
        {
          "id": 2102062,
          "postDate": "2023-01-16T11:09:13.710Z",
          "content": "<p>Yes, it is.<br>\nPlease check cell 10 of <a href=\"https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook\" target=\"_blank\">my notebook</a>.<br>\nIt was surprising for me.</p>",
          "rawMarkdown": "Yes, it is.\nPlease check cell 10 of [my notebook](https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook).\nIt was surprising for me.",
          "votes": 1,
          "replies": [
            {
              "id": 2102066,
              "postDate": "2023-01-16T11:14:59.573Z",
              "content": "<p>thank you.</p>",
              "rawMarkdown": "thank you.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2055387,
      "postDate": "2022-12-05T03:52:18.257Z",
      "content": "<p>all zero submission, LB = ???<br>\nall one submission, LB = ???<br>\nfrom this we can compute number/ratio of pos images  (for public)</p>",
      "rawMarkdown": "all zero submission, LB = ???\nall one submission, LB = ???\nfrom this we can compute number/ratio of pos images  (for public)",
      "votes": 1,
      "replies": [
        {
          "id": 2055398,
          "postDate": "2022-12-05T04:04:31.237Z",
          "content": "<p>All zero submission?  That's just zero, right ?</p>",
          "rawMarkdown": "All zero submission?  That's just zero, right ?",
          "votes": 1
        },
        {
          "id": 2055406,
          "postDate": "2022-12-05T04:10:56.487Z",
          "content": "<p>maybe more correctly:</p>\n<ol>\n<li>we are given 4 test images (verify that these are in public)</li>\n<li>submit the score of 4 test images = 0.5, rest of test images with score = 0.</li>\n<li>submit the score of 4 test images = 0.5, rest of test images with score = 1.</li>\n</ol>\n<p>now we \"know\" the label of the 4 test images (from your models or additional probing), then you can work out number/ratio of pos images (for public)</p>",
          "rawMarkdown": "maybe more correctly:\n1. we are given 4 test images (verify that these are in public)\n2. submit the score of 4 test images = 0.5, rest of test images with score = 0.\n3. submit the score of 4 test images = 0.5, rest of test images with score = 1.\n\nnow we \"know\" the label of the 4 test images (from your models or additional probing), then you can work out number/ratio of pos images (for public)\n",
          "votes": 4
        },
        {
          "id": 2055409,
          "postDate": "2022-12-05T04:15:59.743Z",
          "content": "<p>You'd have to do some shenanigans with your personal leaderboard I think to zero in on an accurate number with only two decimals to work with.  Or maybe there's some way to do this in parts.  Haven't really thought about it.</p>\n<p>I think it's worth doing though, will look at it tomorrow.  The incidence should be roughly around .02 or the split was messed with.</p>",
          "rawMarkdown": "You'd have to do some shenanigans with your personal leaderboard I think to zero in on an accurate number with only two decimals to work with.  Or maybe there's some way to do this in parts.  Haven't really thought about it.\n\nI think it's worth doing though, will look at it tomorrow.  The incidence should be roughly around .02 or the split was messed with.",
          "votes": 1
        },
        {
          "id": 2055567,
          "postDate": "2022-12-05T07:09:54.360Z",
          "content": "<p>submitting all ones is enough btw</p>\n<p>follow the formula </p>\n<p><code>mean_cancer = 0.5*LB_score / (1- 0.5*LB_score)</code> </p>\n<p>which follows from </p>\n<p><code>F1 = 2TP / (2TP + FN + FP)</code></p>\n<p>when using that FN = 0 and FP = 1-TP in a submission of all ones</p>",
          "rawMarkdown": "submitting all ones is enough btw\n\nfollow the formula \n\n`mean_cancer = 0.5*LB_score / (1- 0.5*LB_score)` \n\nwhich follows from \n\n`F1 = 2TP / (2TP + FN + FP)`\n\nwhen using that FN = 0 and FP = 1-TP in a submission of all ones",
          "votes": 4
        },
        {
          "id": 2055591,
          "postDate": "2022-12-05T07:56:22.283Z",
          "content": "<p>Share my result<br>\nall one LB: 0.04<br>\nmean_cancer: ~0.02</p>",
          "rawMarkdown": "Share my result\nall one LB: 0.04\nmean_cancer: ~0.02\n",
          "votes": 4
        },
        {
          "id": 2055614,
          "postDate": "2022-12-05T08:27:53Z",
          "content": "<p>Thank you for testing!<br>\nConsidering the decimal rounding, LB score is 0.035-0.045, and positive rate would be 0.0178-0.0230.<br>\nThe value is similar to that of train set (0.0206).</p>",
          "rawMarkdown": "Thank you for testing!\nConsidering the decimal rounding, LB score is 0.035-0.045, and positive rate would be 0.0178-0.0230.\nThe value is similar to that of train set (0.0206).",
          "votes": 2
        },
        {
          "id": 2055767,
          "postDate": "2022-12-05T12:43:03.613Z",
          "content": "<p>i am thinking of something like this:</p>\n<p>for each machine id sort the prediction by score. <br>\nthreshold by score value or score ranking?</p>\n<hr>\n<p>a note on threshold by score. binary threshold may not be the best. there are other way to re calibrate the prediction.</p>",
          "rawMarkdown": "i am thinking of something like this:\n\nfor each machine id sort the prediction by score. \nthreshold by score value or score ranking?\n\n---\n\na note on threshold by score. binary threshold may not be the best. there are other way to re calibrate the prediction.",
          "votes": 1
        },
        {
          "id": 2055836,
          "postDate": "2022-12-05T13:35:30.793Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2060722,
          "postDate": "2022-12-10T10:50:22.143Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> LB decimal rounding is 0.4 = [0.0400000; 0.04999999] though</p>",
          "rawMarkdown": "@tomooinubushi LB decimal rounding is 0.4 = [0.0400000; 0.04999999] though",
          "votes": 1
        },
        {
          "id": 2060840,
          "postDate": "2022-12-10T13:54:31.177Z",
          "content": "<p>Thank you for pointing it out! I didn't know that. <br>\nSo, the positive rate is around 0.0204-0.0256, which is still similar to that of train set.</p>",
          "rawMarkdown": "Thank you for pointing it out! I didn't know that. \nSo, the positive rate is around 0.0204-0.0256, which is still similar to that of train set."
        }
      ]
    },
    {
      "id": 2055362,
      "postDate": "2022-12-05T03:22:56.537Z",
      "content": "<p>test_machine_49_count = len(test_df.query(\"machine_id == 49\"))<br>\ntest_len = len(test_df)<br>\ntest_machine_49_count/test_len &gt; 0.40 is True.  (.43 in traindf)</p>",
      "rawMarkdown": "test_machine_49_count = len(test_df.query(\"machine_id == 49\"))\ntest_len = len(test_df)\ntest_machine_49_count/test_len > 0.40 is True.  (.43 in traindf)\n",
      "votes": 1,
      "replies": [
        {
          "id": 2055404,
          "postDate": "2022-12-05T04:10:09.593Z",
          "content": "<p>Thank you for your comment. It was True.<br>\nSo, we don't need to be afraid too much of hidden new machine IDs.</p>",
          "rawMarkdown": "Thank you for your comment. It was True.\nSo, we don't need to be afraid too much of hidden new machine IDs."
        }
      ]
    },
    {
      "id": 2054484,
      "postDate": "2022-12-04T07:32:26.487Z",
      "content": "<p>you can verify this:<br>\nall machine_id has at least some cancer images at hidden test</p>",
      "rawMarkdown": "you can verify this:\nall machine\\_id has at least some cancer images at hidden test",
      "votes": 1,
      "replies": [
        {
          "id": 2054527,
          "postDate": "2022-12-04T08:06:24.417Z",
          "content": "<p>Do you mean all machines have at least one cancer image?<br>\nHow can I test it? Is it based on LB score values?</p>",
          "rawMarkdown": "Do you mean all machines have at least one cancer image?\nHow can I test it? Is it based on LB score values?"
        },
        {
          "id": 2055381,
          "postDate": "2022-12-05T03:44:37.653Z",
          "content": "<p>Yeah, I'm kinda curious how he figured this out as well.   Assuming the distributions are roughly the same, you probably only have to check 190 and 197.   It's possible that machine_id is simply not a signal at all and we've just got a weird split with regards to those.  Not sure why'd they include it in the test data if it wasn't a signal, tbh.  But I'm also confused as to why density isn't in the test set.</p>",
          "rawMarkdown": "Yeah, I'm kinda curious how he figured this out as well.   Assuming the distributions are roughly the same, you probably only have to check 190 and 197.   It's possible that machine_id is simply not a signal at all and we've just got a weird split with regards to those.  Not sure why'd they include it in the test data if it wasn't a signal, tbh.  But I'm also confused as to why density isn't in the test set.\n\n"
        },
        {
          "id": 2055790,
          "postDate": "2022-12-05T12:54:25.550Z",
          "content": "<p>I thought he had some way of doing it without submitting for each machine id, but it occurs to me that there is some value in actually submitting for each one.   It would be interesting to see if the scors line up with the traindf.</p>\n<p>Eg, using train</p>\n<pre><code> k  train[].unique():\n    (k, pfbeta(train[], [  x != k    x  train[]]))\n</code></pre>\n<pre><code> \n \n \n \n \n \n \n \n \n \n</code></pre>\n<p>Be good to verify roughly the count of how many new machine_ids we see as well and there over all score.,</p>",
          "rawMarkdown": "I thought he had some way of doing it without submitting for each machine id, but it occurs to me that there is some value in actually submitting for each one.   It would be interesting to see if the scors line up with the traindf.\n\nEg, using train\n\n```python\nfor k in train['machine_id'].unique():\n    print(k, pfbeta(train['cancer'], [0 if x != k else 1 for x in train['machine_id']]))\n```\n\n```python\n29 0.03416445623342175\n21 0.03283932188932722\n216 0.005218525766470972\n93 0.009111617312072893\n49 0.04974277960059951\n48 0.03631936694734706\n170 0.02210475732820759\n210 0\n190 0.007674597083653109\n197 0\n```\n\nBe good to verify roughly the count of how many new machine_ids we see as well and there over all score.,"
        },
        {
          "id": 2055814,
          "postDate": "2022-12-05T13:10:07.027Z",
          "content": "<p>for 210 0, 197 0, i suspect the (few) pos cases are in the test</p>",
          "rawMarkdown": "for 210 0, 197 0, i suspect the (few) pos cases are in the test"
        },
        {
          "id": 2056309,
          "postDate": "2022-12-06T01:41:50.247Z",
          "content": "<p>Scores don't seem to be correlating for machine_ids, so either incidence is off or distribution of machine usage.</p>",
          "rawMarkdown": "Scores don't seem to be correlating for machine_ids, so either incidence is off or distribution of machine usage."
        }
      ]
    },
    {
      "id": 2054476,
      "postDate": "2022-12-04T07:16:59.677Z",
      "content": "<p>As with train dataset, the average site id was &lt;1.5. I wish they were all 2;)</p>",
      "rawMarkdown": "As with train dataset, the average site id was <1.5. I wish they were all 2;)",
      "votes": 1,
      "replies": [
        {
          "id": 2054524,
          "postDate": "2022-12-04T08:01:24.697Z",
          "content": "<p>Thank you for your comment. I tested it, and it was true. <br>\nDo you think images in site id 2 are easy to diagnose?</p>",
          "rawMarkdown": "Thank you for your comment. I tested it, and it was true. \nDo you think images in site id 2 are easy to diagnose?",
          "votes": 1
        },
        {
          "id": 2054531,
          "postDate": "2022-12-04T08:10:18.697Z",
          "content": "<p>Because the site id's for the given test were all 2.<br>\nI thought I could further increase the score by devising a validation set if they are all 2.</p>",
          "rawMarkdown": "Because the site id's for the given test were all 2.\nI thought I could further increase the score by devising a validation set if they are all 2.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2055424,
      "postDate": "2022-12-05T04:32:58.397Z",
      "content": "<p>Some other distribution things to check, just to verify it was a random split and not something else:</p>\n<p>Is mean age around 57.5 at site 1, 60 at site2?  Maybe verify the std and nancount as well<br>\nAre there about 0.000644 AT views?  (AT views mean a lesion was found, 10% incidence of cancer in train set)<br>\nImplants around 5% (often results in lower incidence of cancer due to traits of women who get implants)</p>",
      "rawMarkdown": "Some other distribution things to check, just to verify it was a random split and not something else:\n\nIs mean age around 57.5 at site 1, 60 at site2?  Maybe verify the std and nancount as well\nAre there about 0.000644 AT views?  (AT views mean a lesion was found, 10% incidence of cancer in train set)\nImplants around 5% (often results in lower incidence of cancer due to traits of women who get implants)\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2055620,
          "postDate": "2022-12-05T08:32:43.597Z",
          "content": "<p>I verified mean age and implant rate. Together with the positive rate from LB score, the split seems to be random.</p>",
          "rawMarkdown": "I verified mean age and implant rate. Together with the positive rate from LB score, the split seems to be random."
        },
        {
          "id": 2055763,
          "postDate": "2022-12-05T12:38:42.587Z",
          "content": "<p>Yeah, looks good!</p>\n<p>I'm going to check out the implants being reported only in site 1 as well.  I like how you're verifying all the natural assumptions one would make.</p>\n<p>Some deeper probing on the machine ids is probably a good idea as that is a surprise.</p>",
          "rawMarkdown": "Yeah, looks good!\n\nI'm going to check out the implants being reported only in site 1 as well.  I like how you're verifying all the natural assumptions one would make.\n\nSome deeper probing on the machine ids is probably a good idea as that is a surprise."
        }
      ]
    },
    {
      "id": 2054654,
      "postDate": "2022-12-04T10:20:59.223Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> </p>\n<ul>\n<li>Patient IDs in train and test sets do not overlap</li>\n<li>Image IDs in train and test sets do not overlap</li>\n</ul>\n<p>That's what I was wondering. Now we know better how to build our CV.</p>",
      "rawMarkdown": "Thanks @tomooinubushi \n\n- Patient IDs in train and test sets do not overlap\n- Image IDs in train and test sets do not overlap\n\nThat's what I was wondering. Now we know better how to build our CV.\n",
      "votes": 2
    },
    {
      "id": 2149352,
      "postDate": "2023-02-18T08:00:03.670Z",
      "content": "<p>Apologies If I am asking a duplicate question…what is the total number of patients in the hidden test set?</p>",
      "rawMarkdown": "Apologies If I am asking a duplicate question...what is the total number of patients in the hidden test set?",
      "replies": [
        {
          "id": 2149399,
          "postDate": "2023-02-18T09:07:44.997Z",
          "content": "<p>\"You can expect roughly 8,000 patients in the hidden test set. \" - from Data section of this competition.</p>",
          "rawMarkdown": "\"You can expect roughly 8,000 patients in the hidden test set. \" - from Data section of this competition.",
          "votes": 2,
          "replies": [
            {
              "id": 2149439,
              "postDate": "2023-02-18T09:51:30.327Z",
              "content": "<p>thanks a lot for clarifying…much appreciated </p>",
              "rawMarkdown": "thanks a lot for clarifying...much appreciated "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2148868,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-17T18:14:19.580000",
      "content": "<blockquote>\n  <p>All patients have CC and MLO images for both sides</p>\n</blockquote>\n<p>This is surprising, as there seemed to be examinations in the <strong>train set</strong> where at least one of the views were missing (e.g., 5 <code>DICOM</code> files, 3 L-CC, 1 L-MLO and 1 R-MLO, but no R-CC).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2101738,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2023-01-16T06:26:37.303000",
      "content": "<p>Thank you. Is this true: '<em>There are new machine IDs in test dataset'</em> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2102062,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2023-01-16T11:09:13.710000",
          "content": "<p>Yes, it is.<br>\nPlease check cell 10 of <a href=\"https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook\" target=\"_blank\">my notebook</a>.<br>\nIt was surprising for me.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2102066,
              "author_name": "GUNER",
              "author_url": "",
              "post_date": "2023-01-16T11:14:59.573000",
              "content": "<p>thank you.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2055387,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-05T03:52:18.257000",
      "content": "<p>all zero submission, LB = ???<br>\nall one submission, LB = ???<br>\nfrom this we can compute number/ratio of pos images  (for public)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2055398,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T04:04:31.237000",
          "content": "<p>All zero submission?  That's just zero, right ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2055406,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-05T04:10:56.487000",
          "content": "<p>maybe more correctly:</p>\n<ol>\n<li>we are given 4 test images (verify that these are in public)</li>\n<li>submit the score of 4 test images = 0.5, rest of test images with score = 0.</li>\n<li>submit the score of 4 test images = 0.5, rest of test images with score = 1.</li>\n</ol>\n<p>now we \"know\" the label of the 4 test images (from your models or additional probing), then you can work out number/ratio of pos images (for public)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2055409,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T04:15:59.743000",
          "content": "<p>You'd have to do some shenanigans with your personal leaderboard I think to zero in on an accurate number with only two decimals to work with.  Or maybe there's some way to do this in parts.  Haven't really thought about it.</p>\n<p>I think it's worth doing though, will look at it tomorrow.  The incidence should be roughly around .02 or the split was messed with.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2055567,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2022-12-05T07:09:54.360000",
          "content": "<p>submitting all ones is enough btw</p>\n<p>follow the formula </p>\n<p><code>mean_cancer = 0.5*LB_score / (1- 0.5*LB_score)</code> </p>\n<p>which follows from </p>\n<p><code>F1 = 2TP / (2TP + FN + FP)</code></p>\n<p>when using that FN = 0 and FP = 1-TP in a submission of all ones</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2055591,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2022-12-05T07:56:22.283000",
          "content": "<p>Share my result<br>\nall one LB: 0.04<br>\nmean_cancer: ~0.02</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2055614,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-05T08:27:53",
          "content": "<p>Thank you for testing!<br>\nConsidering the decimal rounding, LB score is 0.035-0.045, and positive rate would be 0.0178-0.0230.<br>\nThe value is similar to that of train set (0.0206).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2055767,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-05T12:43:03.613000",
          "content": "<p>i am thinking of something like this:</p>\n<p>for each machine id sort the prediction by score. <br>\nthreshold by score value or score ranking?</p>\n<hr>\n<p>a note on threshold by score. binary threshold may not be the best. there are other way to re calibrate the prediction.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2055836,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-05T13:35:30.793000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060722,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2022-12-10T10:50:22.143000",
          "content": "<p><a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> LB decimal rounding is 0.4 = [0.0400000; 0.04999999] though</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2060840,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-10T13:54:31.177000",
          "content": "<p>Thank you for pointing it out! I didn't know that. <br>\nSo, the positive rate is around 0.0204-0.0256, which is still similar to that of train set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2055362,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-05T03:22:56.537000",
      "content": "<p>test_machine_49_count = len(test_df.query(\"machine_id == 49\"))<br>\ntest_len = len(test_df)<br>\ntest_machine_49_count/test_len &gt; 0.40 is True.  (.43 in traindf)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2055404,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-05T04:10:09.593000",
          "content": "<p>Thank you for your comment. It was True.<br>\nSo, we don't need to be afraid too much of hidden new machine IDs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2054484,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-04T07:32:26.487000",
      "content": "<p>you can verify this:<br>\nall machine_id has at least some cancer images at hidden test</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2054527,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-04T08:06:24.417000",
          "content": "<p>Do you mean all machines have at least one cancer image?<br>\nHow can I test it? Is it based on LB score values?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055381,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T03:44:37.653000",
          "content": "<p>Yeah, I'm kinda curious how he figured this out as well.   Assuming the distributions are roughly the same, you probably only have to check 190 and 197.   It's possible that machine_id is simply not a signal at all and we've just got a weird split with regards to those.  Not sure why'd they include it in the test data if it wasn't a signal, tbh.  But I'm also confused as to why density isn't in the test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055790,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T12:54:25.550000",
          "content": "<p>I thought he had some way of doing it without submitting for each machine id, but it occurs to me that there is some value in actually submitting for each one.   It would be interesting to see if the scors line up with the traindf.</p>\n<p>Eg, using train</p>\n<pre><code> k  train[].unique():\n    (k, pfbeta(train[], [  x != k    x  train[]]))\n</code></pre>\n<pre><code> \n \n \n \n \n \n \n \n \n \n</code></pre>\n<p>Be good to verify roughly the count of how many new machine_ids we see as well and there over all score.,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055814,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-05T13:10:07.027000",
          "content": "<p>for 210 0, 197 0, i suspect the (few) pos cases are in the test</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056309,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-06T01:41:50.247000",
          "content": "<p>Scores don't seem to be correlating for machine_ids, so either incidence is off or distribution of machine usage.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2054476,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2022-12-04T07:16:59.677000",
      "content": "<p>As with train dataset, the average site id was &lt;1.5. I wish they were all 2;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2054524,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-04T08:01:24.697000",
          "content": "<p>Thank you for your comment. I tested it, and it was true. <br>\nDo you think images in site id 2 are easy to diagnose?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2054531,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2022-12-04T08:10:18.697000",
          "content": "<p>Because the site id's for the given test were all 2.<br>\nI thought I could further increase the score by devising a validation set if they are all 2.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2055424,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-05T04:32:58.397000",
      "content": "<p>Some other distribution things to check, just to verify it was a random split and not something else:</p>\n<p>Is mean age around 57.5 at site 1, 60 at site2?  Maybe verify the std and nancount as well<br>\nAre there about 0.000644 AT views?  (AT views mean a lesion was found, 10% incidence of cancer in train set)<br>\nImplants around 5% (often results in lower incidence of cancer due to traits of women who get implants)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2055620,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-12-05T08:32:43.597000",
          "content": "<p>I verified mean age and implant rate. Together with the positive rate from LB score, the split seems to be random.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2055763,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-05T12:38:42.587000",
          "content": "<p>Yeah, looks good!</p>\n<p>I'm going to check out the implants being reported only in site 1 as well.  I like how you're verifying all the natural assumptions one would make.</p>\n<p>Some deeper probing on the machine ids is probably a good idea as that is a surprise.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2054654,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2022-12-04T10:20:59.223000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> </p>\n<ul>\n<li>Patient IDs in train and test sets do not overlap</li>\n<li>Image IDs in train and test sets do not overlap</li>\n</ul>\n<p>That's what I was wondering. Now we know better how to build our CV.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2149352,
      "author_name": "sandy1112",
      "author_url": "",
      "post_date": "2023-02-18T08:00:03.670000",
      "content": "<p>Apologies If I am asking a duplicate question…what is the total number of patients in the hidden test set?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2149399,
          "author_name": "Rasoul Mojtahedzadeh",
          "author_url": "",
          "post_date": "2023-02-18T09:07:44.997000",
          "content": "<p>\"You can expect roughly 8,000 patients in the hidden test set. \" - from Data section of this competition.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2149439,
              "author_name": "sandy1112",
              "author_url": "",
              "post_date": "2023-02-18T09:51:30.327000",
              "content": "<p>thanks a lot for clarifying…much appreciated </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2054249": "In [this notebook](https://www.kaggle.com/code/tomooinubushi/some-lb-probing-results-to-share/notebook), I tested some basic assumptions about test dataset.\nI demonstrated that following assumptions are all TRUE.\n\n* There are no new site ID in test dataset.\n* Patient IDs in train and test sets do not overlap\n* Image IDs in train and test sets do not overlap\n* There are no new laterality values in test dataset.\n* There are new machine IDs in test dataset. (This is already raised by @abebe9849 in [here](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369362))\n* There are no new view values in test dataset.\n* No. of images/patient are all >= 4\n* Site ID is always the same for each patient.\n* Age is always the same for each patient.\n* There are no overlap of machine IDs between two sites in test dataset.\n* Some patients underwent mammography with multiple machines.\n\n[updated at 2022-12-04]\n* No. of images in site ID 1 > No. of images in site ID 2. (by @yujiariyasu)\n\n[updated at 2022-12-05]\n* All patients have CC and MLO images for both sides\n* More than 40% of images are from machine ID 49 (43% for train dataset) (by @kaggleqrdl)\n* Mean age of patients is between 56-61 (58.6 for train set), and patients in site 1 are younger than those in site 2. (by @kaggleqrdl)\n* 1-2% of patients use implants (1.4% for train set).(by @kaggleqrdl)\n* **Positive rate is between 0.0204-0.0256**  (0.0206 for train dataset). (shown by @zzy990106 and others)\n\n[updated at 2022-12-15]\n* Age column contains nan, while others do not\n\nI am happy if anyone correct me if I am wrong.\nI am also very happy if anyone share us other assumptions/hypothesis about test dataset.",
    "2148868": ">All patients have CC and MLO images for both sides\n\nThis is surprising, as there seemed to be examinations in the **train set** where at least one of the views were missing (e.g., 5 `DICOM` files, 3 L-CC, 1 L-MLO and 1 R-MLO, but no R-CC).",
    "2101738": "Thank you. Is this true: '*There are new machine IDs in test dataset'* ?",
    "2055387": "all zero submission, LB = ???\nall one submission, LB = ???\nfrom this we can compute number/ratio of pos images  (for public)",
    "2055362": "test_machine_49_count = len(test_df.query(\"machine_id == 49\"))\ntest_len = len(test_df)\ntest_machine_49_count/test_len > 0.40 is True.  (.43 in traindf)\n",
    "2054484": "you can verify this:\nall machine\\_id has at least some cancer images at hidden test",
    "2054476": "As with train dataset, the average site id was <1.5. I wish they were all 2;)",
    "2055424": "Some other distribution things to check, just to verify it was a random split and not something else:\n\nIs mean age around 57.5 at site 1, 60 at site2?  Maybe verify the std and nancount as well\nAre there about 0.000644 AT views?  (AT views mean a lesion was found, 10% incidence of cancer in train set)\nImplants around 5% (often results in lower incidence of cancer due to traits of women who get implants)\n\n",
    "2054654": "Thanks @tomooinubushi \n\n- Patient IDs in train and test sets do not overlap\n- Image IDs in train and test sets do not overlap\n\nThat's what I was wondering. Now we know better how to build our CV.\n",
    "2149352": "Apologies If I am asking a duplicate question...what is the total number of patients in the hidden test set?"
  }
}