{
  "id": 164009,
  "title": "Train 2 different models in 2 centers",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/164009",
  "author_name": "seefun",
  "post_date": "2020-07-04T11:45:43.333000",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>There are two different centers(data_provider) \"Radboud“ and \"Karolinska\".\nIf the data of the new center does not appear in the test set, can we train different models in different centers？\nI try to finetune my model (use both two centers' data) in these two centers separately，I got a gain of about 0.003 QWK in my local valset. (eg. 0.880-&gt;0.883, both \"Radboud“ and \"Karolinska\")</p>",
  "messages": [
    {
      "id": 914964,
      "postDate": "2020-07-04T11:45:43.333Z",
      "content": "<p>There are two different centers(data_provider) \"Radboud“ and \"Karolinska\".\nIf the data of the new center does not appear in the test set, can we train different models in different centers？\nI try to finetune my model (use both two centers' data) in these two centers separately，I got a gain of about 0.003 QWK in my local valset. (eg. 0.880-&gt;0.883, both \"Radboud“ and \"Karolinska\")</p>",
      "rawMarkdown": "There are two different centers(data_provider) \"Radboud“ and \"Karolinska\".\nIf the data of the new center does not appear in the test set, can we train different models in different centers？\nI try to finetune my model (use both two centers' data) in these two centers separately，I got a gain of about 0.003 QWK in my local valset. (eg. 0.880-&gt;0.883, both \"Radboud“ and \"Karolinska\")",
      "votes": 1
    },
    {
      "id": 915712,
      "postDate": "2020-07-05T03:26:12.007Z",
      "content": "<p>I trained two models because the provider information is also available in the test set.\nAs a result, Local QWK improved (radboud: 0.8839, karolinska: 0.9112), but lb score worsened (LB: 0.86).</p>",
      "rawMarkdown": "I trained two models because the provider information is also available in the test set.\nAs a result, Local QWK improved (radboud: 0.8839, karolinska: 0.9112), but lb score worsened (LB: 0.86).",
      "replies": [
        {
          "id": 916483,
          "postDate": "2020-07-05T17:39:04.070Z",
          "content": "<p>Have you been careful with duplicates ? It's a big problem for Radboud.</p>",
          "rawMarkdown": "Have you been careful with duplicates ? It's a big problem for Radboud."
        },
        {
          "id": 916592,
          "postDate": "2020-07-05T20:07:47.127Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> did you remove duplicates or consider them as an augmentation ? Did it boost your score ? </p>",
          "rawMarkdown": "@arroqc did you remove duplicates or consider them as an augmentation ? Did it boost your score ? "
        },
        {
          "id": 916640,
          "postDate": "2020-07-05T21:39:17.860Z",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> Even though I've created a split that takes this issue and other into account I haven't yet started to train with it.</p>",
          "rawMarkdown": "@alexj21 Even though I've created a split that takes this issue and other into account I haven't yet started to train with it."
        },
        {
          "id": 916738,
          "postDate": "2020-07-06T01:27:05.280Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a>  I have removed duplicates using threshold 0.96 based on kernel <a href=\"https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\">here</a>.</p>\n\n<p><a href=\"/alexj21\">@alexj21</a>  I didn't train separately before removing the duplicates, but there was a jump of 0.86-&gt;0.88 LB by removing duplicates in one model training. (1 fold)\nHowever, I can't be sure if it works because LB is so blurry :(</p>",
          "rawMarkdown": "@arroqc  I have removed duplicates using threshold 0.96 based on kernel [here](https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images).\n\n@alexj21  I didn't train separately before removing the duplicates, but there was a jump of 0.86-&gt;0.88 LB by removing duplicates in one model training. (1 fold)\nHowever, I can't be sure if it works because LB is so blurry :("
        },
        {
          "id": 917013,
          "postDate": "2020-07-06T07:11:55.077Z",
          "content": "<p>Thanks for your answers, it's good to know. \n<a href=\"/tattaka\">@tattaka</a> did you observe a lower CV/loss ?</p>",
          "rawMarkdown": "Thanks for your answers, it's good to know. \n@tattaka did you observe a lower CV/loss ?"
        },
        {
          "id": 917056,
          "postDate": "2020-07-06T07:54:22.757Z",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> I got the similar CV compared to before removing the duplicates.</p>",
          "rawMarkdown": "@alexj21 I got the similar CV compared to before removing the duplicates."
        },
        {
          "id": 917061,
          "postDate": "2020-07-06T07:56:05.160Z",
          "content": "<p>interesting ! thanks</p>",
          "rawMarkdown": "interesting ! thanks"
        }
      ]
    },
    {
      "id": 915566,
      "postDate": "2020-07-04T20:48:03.013Z",
      "content": "<p>Are you suggesting to train 2 different models, one for each institute?</p>",
      "rawMarkdown": "Are you suggesting to train 2 different models, one for each institute?",
      "replies": [
        {
          "id": 915657,
          "postDate": "2020-07-05T01:55:03.307Z",
          "content": "<p>I don’t know if this method should be chosen because it increases the risk of overfitting and the improvement is extremely limited</p>",
          "rawMarkdown": "I don’t know if this method should be chosen because it increases the risk of overfitting and the improvement is extremely limited"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 915712,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2020-07-05T03:26:12.007000",
      "content": "<p>I trained two models because the provider information is also available in the test set.\nAs a result, Local QWK improved (radboud: 0.8839, karolinska: 0.9112), but lb score worsened (LB: 0.86).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 916483,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-07-05T17:39:04.070000",
          "content": "<p>Have you been careful with duplicates ? It's a big problem for Radboud.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916592,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-07-05T20:07:47.127000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> did you remove duplicates or consider them as an augmentation ? Did it boost your score ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916640,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-07-05T21:39:17.860000",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> Even though I've created a split that takes this issue and other into account I haven't yet started to train with it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 916738,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-07-06T01:27:05.280000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a>  I have removed duplicates using threshold 0.96 based on kernel <a href=\"https://www.kaggle.com/appian/panda-imagehash-to-detect-duplicate-images\">here</a>.</p>\n\n<p><a href=\"/alexj21\">@alexj21</a>  I didn't train separately before removing the duplicates, but there was a jump of 0.86-&gt;0.88 LB by removing duplicates in one model training. (1 fold)\nHowever, I can't be sure if it works because LB is so blurry :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917013,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-07-06T07:11:55.077000",
          "content": "<p>Thanks for your answers, it's good to know. \n<a href=\"/tattaka\">@tattaka</a> did you observe a lower CV/loss ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917056,
          "author_name": "tattaka",
          "author_url": "",
          "post_date": "2020-07-06T07:54:22.757000",
          "content": "<p><a href=\"/alexj21\">@alexj21</a> I got the similar CV compared to before removing the duplicates.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917061,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-07-06T07:56:05.160000",
          "content": "<p>interesting ! thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915566,
      "author_name": "Pasquale",
      "author_url": "",
      "post_date": "2020-07-04T20:48:03.013000",
      "content": "<p>Are you suggesting to train 2 different models, one for each institute?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 915657,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2020-07-05T01:55:03.307000",
          "content": "<p>I don’t know if this method should be chosen because it increases the risk of overfitting and the improvement is extremely limited</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "914964": "There are two different centers(data_provider) \"Radboud“ and \"Karolinska\".\nIf the data of the new center does not appear in the test set, can we train different models in different centers？\nI try to finetune my model (use both two centers' data) in these two centers separately，I got a gain of about 0.003 QWK in my local valset. (eg. 0.880-&gt;0.883, both \"Radboud“ and \"Karolinska\")",
    "915712": "I trained two models because the provider information is also available in the test set.\nAs a result, Local QWK improved (radboud: 0.8839, karolinska: 0.9112), but lb score worsened (LB: 0.86).",
    "915566": "Are you suggesting to train 2 different models, one for each institute?"
  }
}