{
  "id": 169212,
  "title": "What worked and what didn't work",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/169212",
  "author_name": "Kirderf",
  "post_date": "2020-07-23T08:23:25.487000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks for an important task and competition, and as a swede it's extra fun to see Karolinska as a sponsor at Kaggle, keep it up!</p>\n\n<p>First and foremost, congrats to the winners, great work! :)</p>\n\n<p>108th place isn't to cheer for but I think I have learned of few things, which is the important factor. I got that score with only one single efficientnet-b0 model(credit <a href=\"/haqishen\">@haqishen</a>), used it as a benchmark to beat and didn't because of lower resolution I think, so with more training, diff models and ensembles of folds the score should have been better.</p>\n\n<p>What didn't work:\n- I got stuck in solving the large image dimensions, tried several of solutions, the training and tuning added to the different image dim testing, took to much time, biggest failure for me.\n- AdamW( plain Adam took the lead after 5 epochs).\n- 1Cycle(w/ Adam, AdamW), ReduceLROnPlateau, Ranger, RangerLars, SWA.\n- Left with 10min GPU-time the last day for submitting, not a big factor this time but too stressfull and resulted in that I couldn't ensemble the last solutions, my bad, and cost some places.\n- efficientnet-b0 worked better than efficientnet-b0 noisy student.</p>\n\n<p>What did work:\n- Full resolution and dimensions and with tiles, everything worked much better! But time was a factor so didn't have time to create new models, open the last \"testing-box\" to late.\n- Adam w/ the public solution of CosineAnnealingLR and GradualWarmupScheduler(w/ the version fix it worked better). Great schedule mix, I will try it again in next competitions.\n- SGD w/ 0.9 momentum and nesterov + 1Cycle, like orginal research paper but with total steps ( 6 epochs x steps per epochs) and more deeper lr 1e-3 + 1e-6. Worked well!\n- OptimizedRounder with input from training/validation.\n- Predict/inference using single model with 14 TTA / 14 ensemble. The accuracy boost went up for every added TTA, I stoped with 14 but had tried bigger TTA with more time.</p>\n\n<p>Thank you all, great teamwork and contributions!</p>",
  "messages": [
    {
      "id": 941464,
      "postDate": "2020-07-23T08:23:25.487Z",
      "content": "<p>Thanks for an important task and competition, and as a swede it's extra fun to see Karolinska as a sponsor at Kaggle, keep it up!</p>\n\n<p>First and foremost, congrats to the winners, great work! :)</p>\n\n<p>108th place isn't to cheer for but I think I have learned of few things, which is the important factor. I got that score with only one single efficientnet-b0 model(credit <a href=\"/haqishen\">@haqishen</a>), used it as a benchmark to beat and didn't because of lower resolution I think, so with more training, diff models and ensembles of folds the score should have been better.</p>\n\n<p>What didn't work:\n- I got stuck in solving the large image dimensions, tried several of solutions, the training and tuning added to the different image dim testing, took to much time, biggest failure for me.\n- AdamW( plain Adam took the lead after 5 epochs).\n- 1Cycle(w/ Adam, AdamW), ReduceLROnPlateau, Ranger, RangerLars, SWA.\n- Left with 10min GPU-time the last day for submitting, not a big factor this time but too stressfull and resulted in that I couldn't ensemble the last solutions, my bad, and cost some places.\n- efficientnet-b0 worked better than efficientnet-b0 noisy student.</p>\n\n<p>What did work:\n- Full resolution and dimensions and with tiles, everything worked much better! But time was a factor so didn't have time to create new models, open the last \"testing-box\" to late.\n- Adam w/ the public solution of CosineAnnealingLR and GradualWarmupScheduler(w/ the version fix it worked better). Great schedule mix, I will try it again in next competitions.\n- SGD w/ 0.9 momentum and nesterov + 1Cycle, like orginal research paper but with total steps ( 6 epochs x steps per epochs) and more deeper lr 1e-3 + 1e-6. Worked well!\n- OptimizedRounder with input from training/validation.\n- Predict/inference using single model with 14 TTA / 14 ensemble. The accuracy boost went up for every added TTA, I stoped with 14 but had tried bigger TTA with more time.</p>\n\n<p>Thank you all, great teamwork and contributions!</p>",
      "rawMarkdown": "Thanks for an important task and competition, and as a swede it's extra fun to see Karolinska as a sponsor at Kaggle, keep it up!\n\nFirst and foremost, congrats to the winners, great work! :)\n\n108th place isn't to cheer for but I think I have learned of few things, which is the important factor. I got that score with only one single efficientnet-b0 model(credit @haqishen), used it as a benchmark to beat and didn't because of lower resolution I think, so with more training, diff models and ensembles of folds the score should have been better.\n\nWhat didn't work:\n- I got stuck in solving the large image dimensions, tried several of solutions, the training and tuning added to the different image dim testing, took to much time, biggest failure for me.\n- AdamW( plain Adam took the lead after 5 epochs).\n- 1Cycle(w/ Adam, AdamW), ReduceLROnPlateau, Ranger, RangerLars, SWA.\n- Left with 10min GPU-time the last day for submitting, not a big factor this time but too stressfull and resulted in that I couldn't ensemble the last solutions, my bad, and cost some places.\n- efficientnet-b0 worked better than efficientnet-b0 noisy student.\n\nWhat did work:\n- Full resolution and dimensions and with tiles, everything worked much better! But time was a factor so didn't have time to create new models, open the last \"testing-box\" to late.\n- Adam w/ the public solution of CosineAnnealingLR and GradualWarmupScheduler(w/ the version fix it worked better). Great schedule mix, I will try it again in next competitions.\n- SGD w/ 0.9 momentum and nesterov + 1Cycle, like orginal research paper but with total steps ( 6 epochs x steps per epochs) and more deeper lr 1e-3 + 1e-6. Worked well!\n- OptimizedRounder with input from training/validation.\n- Predict/inference using single model with 14 TTA / 14 ensemble. The accuracy boost went up for every added TTA, I stoped with 14 but had tried bigger TTA with more time.\n\nThank you all, great teamwork and contributions!"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "941464": "Thanks for an important task and competition, and as a swede it's extra fun to see Karolinska as a sponsor at Kaggle, keep it up!\n\nFirst and foremost, congrats to the winners, great work! :)\n\n108th place isn't to cheer for but I think I have learned of few things, which is the important factor. I got that score with only one single efficientnet-b0 model(credit @haqishen), used it as a benchmark to beat and didn't because of lower resolution I think, so with more training, diff models and ensembles of folds the score should have been better.\n\nWhat didn't work:\n- I got stuck in solving the large image dimensions, tried several of solutions, the training and tuning added to the different image dim testing, took to much time, biggest failure for me.\n- AdamW( plain Adam took the lead after 5 epochs).\n- 1Cycle(w/ Adam, AdamW), ReduceLROnPlateau, Ranger, RangerLars, SWA.\n- Left with 10min GPU-time the last day for submitting, not a big factor this time but too stressfull and resulted in that I couldn't ensemble the last solutions, my bad, and cost some places.\n- efficientnet-b0 worked better than efficientnet-b0 noisy student.\n\nWhat did work:\n- Full resolution and dimensions and with tiles, everything worked much better! But time was a factor so didn't have time to create new models, open the last \"testing-box\" to late.\n- Adam w/ the public solution of CosineAnnealingLR and GradualWarmupScheduler(w/ the version fix it worked better). Great schedule mix, I will try it again in next competitions.\n- SGD w/ 0.9 momentum and nesterov + 1Cycle, like orginal research paper but with total steps ( 6 epochs x steps per epochs) and more deeper lr 1e-3 + 1e-6. Worked well!\n- OptimizedRounder with input from training/validation.\n- Predict/inference using single model with 14 TTA / 14 ensemble. The accuracy boost went up for every added TTA, I stoped with 14 but had tried bigger TTA with more time.\n\nThank you all, great teamwork and contributions!"
  }
}