{
  "id": 347144,
  "title": "Usage of Supplementary Data during Model Training ",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/347144",
  "author_name": "Syed Muhammad Atif",
  "post_date": "2022-08-23T04:43:18.618000",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Everyone,</p>\n<p>Apart from \"train\" folder, there is an \"other\" folder in data which contain slide that have clots other than CA and LLA. My question is that how a model should be train i.e.,<br>\n1) For a given slide image, it will just classify it as CA or ALL i.e., the model will have two outputs.<br>\n2) For a given slide image, it will just classify it as CA, ALL and other i.e., the model will have three outputs.</p>\n<p>In the former case, the model does not need data in \"other\" folder. However, it is required in the later case and obviously the resulting model is more generalized.<br>\nOtherwise, the data in \"other\" folder is of no need.</p>\n<p>What do you think?</p>",
  "messages": [
    {
      "id": 1912311,
      "postDate": "2022-08-24T16:33:53.337Z",
      "content": "<p>Use another predictor to determine if it need it</p>",
      "rawMarkdown": "Use another predictor to determine if it need it",
      "votes": 1
    },
    {
      "id": 1909960,
      "postDate": "2022-08-23T04:43:18.620Z",
      "content": "<p>Hello Everyone,</p>\n<p>Apart from \"train\" folder, there is an \"other\" folder in data which contain slide that have clots other than CA and LLA. My question is that how a model should be train i.e.,<br>\n1) For a given slide image, it will just classify it as CA or ALL i.e., the model will have two outputs.<br>\n2) For a given slide image, it will just classify it as CA, ALL and other i.e., the model will have three outputs.</p>\n<p>In the former case, the model does not need data in \"other\" folder. However, it is required in the later case and obviously the resulting model is more generalized.<br>\nOtherwise, the data in \"other\" folder is of no need.</p>\n<p>What do you think?</p>",
      "rawMarkdown": "Hello Everyone,\n\nApart from \"train\" folder, there is an \"other\" folder in data which contain slide that have clots other than CA and LLA. My question is that how a model should be train i.e.,\n1) For a given slide image, it will just classify it as CA or ALL i.e., the model will have two outputs.\n2) For a given slide image, it will just classify it as CA, ALL and other i.e., the model will have three outputs.\n\nIn the former case, the model does not need data in \"other\" folder. However, it is required in the later case and obviously the resulting model is more generalized.\nOtherwise, the data in \"other\" folder is of no need.\n\nWhat do you think?",
      "votes": 2
    },
    {
      "id": 1914281,
      "postDate": "2022-08-25T23:59:42.700Z",
      "content": "<p>I guess there could be a few reasons to still consider this 'Other' data:</p>\n<p>1) In case you are working on a MIL type solution, the 'Other' label images can be used to train negative tile instance for your feature extractor and classifier - they definitely do not belong to the two classes we want.</p>\n<p>2) You could also use it in the CNN training, although I don't know how appropriate this method would be - consider that if LAA class is [1,0] and CE class is [0,1] in your targets, 'Other' class would be a 'definitely negative' [0,0] target label. This might help in avoiding overfitting, but must be used carefully in the training process. I have tried this second method myself, it did help to some extent.</p>\n<p>PS. making 'Other' images a separate class would be detrimental to the network training in my opinion. Finally, 'Unknown' label is most confusing and should be treated with utmost care and perhaps only used in some randomized fashion, if at all.</p>",
      "rawMarkdown": "I guess there could be a few reasons to still consider this 'Other' data:\n\n1) In case you are working on a MIL type solution, the 'Other' label images can be used to train negative tile instance for your feature extractor and classifier - they definitely do not belong to the two classes we want.\n\n2) You could also use it in the CNN training, although I don't know how appropriate this method would be - consider that if LAA class is [1,0] and CE class is [0,1] in your targets, 'Other' class would be a 'definitely negative' [0,0] target label. This might help in avoiding overfitting, but must be used carefully in the training process. I have tried this second method myself, it did help to some extent.\n\nPS. making 'Other' images a separate class would be detrimental to the network training in my opinion. Finally, 'Unknown' label is most confusing and should be treated with utmost care and perhaps only used in some randomized fashion, if at all."
    },
    {
      "id": 1913036,
      "postDate": "2022-08-25T05:24:28.987Z",
      "content": "<p><a href=\"https://www.kaggle.com/zhehaoliang\" target=\"_blank\">@zhehaoliang</a> Actually, I am more concern why the competition organizers not clearly elaborate the problem.<br>\nSuppose a participant submit the classifier that works for only two classes CE, LAA but the organizers test samples also have samples belonging to 'other' class (the etiology of the blood clot is either Unknown or Other) then obviously the results of his/her classifier will be poor.</p>",
      "rawMarkdown": "@zhehaoliang Actually, I am more concern why the competition organizers not clearly elaborate the problem.\nSuppose a participant submit the classifier that works for only two classes CE, LAA but the organizers test samples also have samples belonging to 'other' class (the etiology of the blood clot is either Unknown or Other) then obviously the results of his/her classifier will be poor."
    }
  ],
  "comments": [
    {
      "id": 1912311,
      "author_name": "zhehao liang",
      "author_url": "",
      "post_date": "2022-08-24T16:33:53.337000",
      "content": "<p>Use another predictor to determine if it need it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914281,
      "author_name": "tdiceman",
      "author_url": "",
      "post_date": "2022-08-25T23:59:42.700000",
      "content": "<p>I guess there could be a few reasons to still consider this 'Other' data:</p>\n<p>1) In case you are working on a MIL type solution, the 'Other' label images can be used to train negative tile instance for your feature extractor and classifier - they definitely do not belong to the two classes we want.</p>\n<p>2) You could also use it in the CNN training, although I don't know how appropriate this method would be - consider that if LAA class is [1,0] and CE class is [0,1] in your targets, 'Other' class would be a 'definitely negative' [0,0] target label. This might help in avoiding overfitting, but must be used carefully in the training process. I have tried this second method myself, it did help to some extent.</p>\n<p>PS. making 'Other' images a separate class would be detrimental to the network training in my opinion. Finally, 'Unknown' label is most confusing and should be treated with utmost care and perhaps only used in some randomized fashion, if at all.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913036,
      "author_name": "Syed Muhammad Atif",
      "author_url": "",
      "post_date": "2022-08-25T05:24:28.987000",
      "content": "<p><a href=\"https://www.kaggle.com/zhehaoliang\" target=\"_blank\">@zhehaoliang</a> Actually, I am more concern why the competition organizers not clearly elaborate the problem.<br>\nSuppose a participant submit the classifier that works for only two classes CE, LAA but the organizers test samples also have samples belonging to 'other' class (the etiology of the blood clot is either Unknown or Other) then obviously the results of his/her classifier will be poor.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912311": "Use another predictor to determine if it need it",
    "1909960": "Hello Everyone,\n\nApart from \"train\" folder, there is an \"other\" folder in data which contain slide that have clots other than CA and LLA. My question is that how a model should be train i.e.,\n1) For a given slide image, it will just classify it as CA or ALL i.e., the model will have two outputs.\n2) For a given slide image, it will just classify it as CA, ALL and other i.e., the model will have three outputs.\n\nIn the former case, the model does not need data in \"other\" folder. However, it is required in the later case and obviously the resulting model is more generalized.\nOtherwise, the data in \"other\" folder is of no need.\n\nWhat do you think?",
    "1914281": "I guess there could be a few reasons to still consider this 'Other' data:\n\n1) In case you are working on a MIL type solution, the 'Other' label images can be used to train negative tile instance for your feature extractor and classifier - they definitely do not belong to the two classes we want.\n\n2) You could also use it in the CNN training, although I don't know how appropriate this method would be - consider that if LAA class is [1,0] and CE class is [0,1] in your targets, 'Other' class would be a 'definitely negative' [0,0] target label. This might help in avoiding overfitting, but must be used carefully in the training process. I have tried this second method myself, it did help to some extent.\n\nPS. making 'Other' images a separate class would be detrimental to the network training in my opinion. Finally, 'Unknown' label is most confusing and should be treated with utmost care and perhaps only used in some randomized fashion, if at all.",
    "1913036": "@zhehaoliang Actually, I am more concern why the competition organizers not clearly elaborate the problem.\nSuppose a participant submit the classifier that works for only two classes CE, LAA but the organizers test samples also have samples belonging to 'other' class (the etiology of the blood clot is either Unknown or Other) then obviously the results of his/her classifier will be poor."
  }
}