{
  "id": 381073,
  "title": "An interesting found with machine_id",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/381073",
  "author_name": "Chenglu",
  "post_date": "2023-01-25T03:28:44.663000",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We know that the pf1 score of the 2 site_id is different, but actually, the diffs are coming from machine_id. The number of samples of different machine_id varies a lot and so is the score.</p>\n<p>When I was doing some tablet feature experiments, I found that if I add the machine_id as one of the categorical features, the CV score can boost by about 0.04. Since there are new machine_id in the test set, or maybe all of machine_id in the test set are different from the training set, I can not use my method to do predictions, but this at least tells us there are some connections of samples from the same machine_id.</p>\n<p>So I wonder if we could utilize the <code>machine_id</code>?</p>",
  "messages": [
    {
      "id": 2114507,
      "postDate": "2023-01-25T03:28:44.663Z",
      "content": "<p>We know that the pf1 score of the 2 site_id is different, but actually, the diffs are coming from machine_id. The number of samples of different machine_id varies a lot and so is the score.</p>\n<p>When I was doing some tablet feature experiments, I found that if I add the machine_id as one of the categorical features, the CV score can boost by about 0.04. Since there are new machine_id in the test set, or maybe all of machine_id in the test set are different from the training set, I can not use my method to do predictions, but this at least tells us there are some connections of samples from the same machine_id.</p>\n<p>So I wonder if we could utilize the <code>machine_id</code>?</p>",
      "rawMarkdown": "We know that the pf1 score of the 2 site_id is different, but actually, the diffs are coming from machine_id. The number of samples of different machine_id varies a lot and so is the score.\n\nWhen I was doing some tablet feature experiments, I found that if I add the machine_id as one of the categorical features, the CV score can boost by about 0.04. Since there are new machine_id in the test set, or maybe all of machine_id in the test set are different from the training set, I can not use my method to do predictions, but this at least tells us there are some connections of samples from the same machine_id.\n\nSo I wonder if we could utilize the `machine_id`?\n",
      "votes": 8
    },
    {
      "id": 2114586,
      "postDate": "2023-01-25T05:10:59.473Z",
      "content": "<p>on a side note, the dicom images may not be orginal image. they could be pre-processed.<br>\n(e.g. if you want to prepare  images for machine learning, you usually do some processing to tidy the data)</p>\n<p>the pre-processing may be different for different site and different machine id</p>",
      "rawMarkdown": "on a side note, the dicom images may not be orginal image. they could be pre-processed.\n(e.g. if you want to prepare  images for machine learning, you usually do some processing to tidy the data)\n\nthe pre-processing may be different for different site and different machine id",
      "votes": 2
    },
    {
      "id": 2117993,
      "postDate": "2023-01-27T17:37:25.187Z",
      "content": "<p>I instead opted to use machine_id for custom scaling on image pre-processing. The distribution is fairly poor for it to be used as a feature.</p>",
      "rawMarkdown": "I instead opted to use machine_id for custom scaling on image pre-processing. The distribution is fairly poor for it to be used as a feature."
    },
    {
      "id": 2114582,
      "postDate": "2023-01-25T05:03:18.037Z",
      "content": "<p>\"I can not use my method to do predictions\"</p>\n<p>you can. just add an unknown id. its train sample can be any train image or external data image.</p>\n<p>the boost probably comes from prior:<br>\np (y|x, machine id)</p>\n<p>if you check the the prior disturbution, some machine are more likely to be negative (there are some id without pos cases in train)</p>",
      "rawMarkdown": "\"I can not use my method to do predictions\"\n\nyou can. just add an unknown id. its train sample can be any train image or external data image.\n\nthe boost probably comes from prior:\np (y|x, machine id)\n\nif you check the the prior disturbution, some machine are more likely to be negative (there are some id without pos cases in train)",
      "replies": [
        {
          "id": 2114621,
          "postDate": "2023-01-25T05:55:40.943Z",
          "content": "<p>\"the boost probably comes from prior\" - I think that's true. </p>\n<p>I have been thinking about adding an unknown id, but I just think it may not be help, to use the machine id , we may have to rely on the model predictions, and, I don't know, maybe some clustering tech?</p>",
          "rawMarkdown": "\"the boost probably comes from prior\" - I think that's true. \n\nI have been thinking about adding an unknown id, but I just think it may not be help, to use the machine id , we may have to rely on the model predictions, and, I don't know, maybe some clustering tech?",
          "replies": [
            {
              "id": 2114632,
              "postDate": "2023-01-25T06:09:19.417Z",
              "content": "<p>just use arcface and machine id feature = [ dist from id 1,  dist from id 2,  dist from id 3,  …]<br>\nor train a softmax classifier and use logit as feature</p>\n<h1>---</h1>\n<p>but for me i would prove that it would work first</p>\n<pre><code>old results\nLB xxx\nprediction = old_model(x)\n\n\nnew results\n\nprediction1 = old_model(x)\nprediction2 = machine_id_trained_model(x  subset, those with train machine id only)\n\nprediction = prediction1 if machine id not found else prediction2\ncheck if this LB has better results\n</code></pre>",
              "rawMarkdown": "just use arcface and machine id feature = [ dist from id 1,  dist from id 2,  dist from id 3,  ...]\nor train a softmax classifier and use logit as feature\n\n#---\nbut for me i would prove that it would work first\n\n```\nold results\nLB xxx\nprediction = old_model(x)\n\n\nnew results\n\nprediction1 = old_model(x)\nprediction2 = machine_id_trained_model(x  subset, those with train machine id only)\n\nprediction = prediction1 if machine id not found else prediction2\ncheck if this LB has better results\n```",
              "votes": 5
            },
            {
              "id": 2118863,
              "postDate": "2023-01-28T10:31:38.413Z",
              "content": "<p>example of gated solution<br>\n<img src=\"https://i.ibb.co/582LpF5/Selection-718.png\" alt=\"https://i.ibb.co/582LpF5/Selection-718.png\"></p>",
              "rawMarkdown": "example of gated solution\n![https://i.ibb.co/582LpF5/Selection-718.png](https://i.ibb.co/582LpF5/Selection-718.png)",
              "votes": 6
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2114586,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-25T05:10:59.473000",
      "content": "<p>on a side note, the dicom images may not be orginal image. they could be pre-processed.<br>\n(e.g. if you want to prepare  images for machine learning, you usually do some processing to tidy the data)</p>\n<p>the pre-processing may be different for different site and different machine id</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2117993,
      "author_name": "AbdouE",
      "author_url": "",
      "post_date": "2023-01-27T17:37:25.187000",
      "content": "<p>I instead opted to use machine_id for custom scaling on image pre-processing. The distribution is fairly poor for it to be used as a feature.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2114582,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-01-25T05:03:18.037000",
      "content": "<p>\"I can not use my method to do predictions\"</p>\n<p>you can. just add an unknown id. its train sample can be any train image or external data image.</p>\n<p>the boost probably comes from prior:<br>\np (y|x, machine id)</p>\n<p>if you check the the prior disturbution, some machine are more likely to be negative (there are some id without pos cases in train)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2114621,
          "author_name": "Chenglu",
          "author_url": "",
          "post_date": "2023-01-25T05:55:40.943000",
          "content": "<p>\"the boost probably comes from prior\" - I think that's true. </p>\n<p>I have been thinking about adding an unknown id, but I just think it may not be help, to use the machine id , we may have to rely on the model predictions, and, I don't know, maybe some clustering tech?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2114632,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-25T06:09:19.417000",
              "content": "<p>just use arcface and machine id feature = [ dist from id 1,  dist from id 2,  dist from id 3,  …]<br>\nor train a softmax classifier and use logit as feature</p>\n<h1>---</h1>\n<p>but for me i would prove that it would work first</p>\n<pre><code>old results\nLB xxx\nprediction = old_model(x)\n\n\nnew results\n\nprediction1 = old_model(x)\nprediction2 = machine_id_trained_model(x  subset, those with train machine id only)\n\nprediction = prediction1 if machine id not found else prediction2\ncheck if this LB has better results\n</code></pre>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2118863,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-28T10:31:38.413000",
              "content": "<p>example of gated solution<br>\n<img src=\"https://i.ibb.co/582LpF5/Selection-718.png\" alt=\"https://i.ibb.co/582LpF5/Selection-718.png\"></p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2114507": "We know that the pf1 score of the 2 site_id is different, but actually, the diffs are coming from machine_id. The number of samples of different machine_id varies a lot and so is the score.\n\nWhen I was doing some tablet feature experiments, I found that if I add the machine_id as one of the categorical features, the CV score can boost by about 0.04. Since there are new machine_id in the test set, or maybe all of machine_id in the test set are different from the training set, I can not use my method to do predictions, but this at least tells us there are some connections of samples from the same machine_id.\n\nSo I wonder if we could utilize the `machine_id`?\n",
    "2114586": "on a side note, the dicom images may not be orginal image. they could be pre-processed.\n(e.g. if you want to prepare  images for machine learning, you usually do some processing to tidy the data)\n\nthe pre-processing may be different for different site and different machine id",
    "2117993": "I instead opted to use machine_id for custom scaling on image pre-processing. The distribution is fairly poor for it to be used as a feature.",
    "2114582": "\"I can not use my method to do predictions\"\n\nyou can. just add an unknown id. its train sample can be any train image or external data image.\n\nthe boost probably comes from prior:\np (y|x, machine id)\n\nif you check the the prior disturbution, some machine are more likely to be negative (there are some id without pos cases in train)"
  }
}