{
  "id": 168825,
  "title": "Which backbone is the best?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/168825",
  "author_name": "seefun",
  "post_date": "2020-07-22T03:27:09.166000",
  "votes": 0,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have tried two ResNet-50 level backbone (ResNet50-gn &amp; tresnet-m) with smooth L1 regression, I can only get 0.88LB but 0.91CV. (TTA, k folds ensemble and model blending are not able to give me great improvement in LB, but CV increased)<br>\nLimited by computing resources, I can’t run more experiments, each experiment takes me a week.<br>\nDoes the choice of backbone matter? Welcome everyone to share the experimental results.</p>",
  "messages": [
    {
      "id": 939286,
      "postDate": "2020-07-22T06:24:21.243Z",
      "content": "<ol>\n<li>efn-b0 here, also limited by computing resources. 5-folds splited but I only choose 2 of 5 folds(2 hightest local score fold) and average them can get 0.888 lb(no tta).</li>\n<li>I try re-sampling the data to train a b0 these days but didn't get improvement.(actually it hurt local score and lb)</li>\n<li>because there are 2 data-providers, I also try feed full-img(after resize, no tiles) to model(a efn-b4) and training with reverse-gradient(use a extra head with 1 output to do binary classification), want to retard the difference between them, but the loss of binary-head become bery huge(can reach 10^2 magnitude), so I give up this way.</li>\n</ol>\n<hr>\n<p>No more experiments, hope my average-model can bring me the good luck. wanna shake 亿点点 too😜</p>",
      "rawMarkdown": "1. efn-b0 here, also limited by computing resources. 5-folds splited but I only choose 2 of 5 folds(2 hightest local score fold) and average them can get 0.888 lb(no tta).\n2. I try re-sampling the data to train a b0 these days but didn't get improvement.(actually it hurt local score and lb)\n3. because there are 2 data-providers, I also try feed full-img(after resize, no tiles) to model(a efn-b4) and training with reverse-gradient(use a extra head with 1 output to do binary classification), want to retard the difference between them, but the loss of binary-head become bery huge(can reach 10^2 magnitude), so I give up this way.\n***\nNo more experiments, hope my average-model can bring me the good luck. wanna shake 亿点点 too😜",
      "votes": 2
    },
    {
      "id": 939109,
      "postDate": "2020-07-22T03:27:09.167Z",
      "content": "<p>I have tried two ResNet-50 level backbone (ResNet50-gn &amp; tresnet-m) with smooth L1 regression, I can only get 0.88LB but 0.91CV. (TTA, k folds ensemble and model blending are not able to give me great improvement in LB, but CV increased)<br>\nLimited by computing resources, I can’t run more experiments, each experiment takes me a week.<br>\nDoes the choice of backbone matter? Welcome everyone to share the experimental results.</p>",
      "rawMarkdown": "I have tried two ResNet-50 level backbone (ResNet50-gn &amp; tresnet-m) with smooth L1 regression, I can only get 0.88LB but 0.91CV. (TTA, k folds ensemble and model blending are not able to give me great improvement in LB, but CV increased)\nLimited by computing resources, I can’t run more experiments, each experiment takes me a week.\nDoes the choice of backbone matter? Welcome everyone to share the experimental results."
    }
  ],
  "comments": [
    {
      "id": 939286,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-07-22T06:24:21.243000",
      "content": "<ol>\n<li>efn-b0 here, also limited by computing resources. 5-folds splited but I only choose 2 of 5 folds(2 hightest local score fold) and average them can get 0.888 lb(no tta).</li>\n<li>I try re-sampling the data to train a b0 these days but didn't get improvement.(actually it hurt local score and lb)</li>\n<li>because there are 2 data-providers, I also try feed full-img(after resize, no tiles) to model(a efn-b4) and training with reverse-gradient(use a extra head with 1 output to do binary classification), want to retard the difference between them, but the loss of binary-head become bery huge(can reach 10^2 magnitude), so I give up this way.</li>\n</ol>\n<hr>\n<p>No more experiments, hope my average-model can bring me the good luck. wanna shake 亿点点 too😜</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "939286": "1. efn-b0 here, also limited by computing resources. 5-folds splited but I only choose 2 of 5 folds(2 hightest local score fold) and average them can get 0.888 lb(no tta).\n2. I try re-sampling the data to train a b0 these days but didn't get improvement.(actually it hurt local score and lb)\n3. because there are 2 data-providers, I also try feed full-img(after resize, no tiles) to model(a efn-b4) and training with reverse-gradient(use a extra head with 1 output to do binary classification), want to retard the difference between them, but the loss of binary-head become bery huge(can reach 10^2 magnitude), so I give up this way.\n***\nNo more experiments, hope my average-model can bring me the good luck. wanna shake 亿点点 too😜",
    "939109": "I have tried two ResNet-50 level backbone (ResNet50-gn &amp; tresnet-m) with smooth L1 regression, I can only get 0.88LB but 0.91CV. (TTA, k folds ensemble and model blending are not able to give me great improvement in LB, but CV increased)\nLimited by computing resources, I can’t run more experiments, each experiment takes me a week.\nDoes the choice of backbone matter? Welcome everyone to share the experimental results."
  }
}