{
  "id": 370689,
  "title": "[LB: 0.37] Tensorflow Baseline with TPU-1VM",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/370689",
  "author_name": "Awsaf",
  "post_date": "2022-12-05T23:27:07.289000",
  "votes": 34,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hi everyone, there has been a new addition to <code>accelerator</code> in Kaggle <a href=\"https://www.kaggle.com/discussions/product-feedback/369338\" target=\"_blank\">here</a>, <code>TPU-1VM</code> aka <code>local TPU</code>. Unlike its predecessor <code>remote-TPU</code>,</p>\n<ul>\n<li>It doesn't require <code>GCS_PATH</code> anymore and it doesn't need <code>internet access</code>, which makes it even more easier to use. </li>\n<li>From the performance perspective unlike <code>remote-TPU</code> its first epoch is not that slow.</li>\n</ul>\n<p>I have published a notebook to show how to train in <code>TPU-1VM</code>, this notebook also supports <code>multi-GPU</code> training.</p>\n<h2>Notebooks</h2>\n<ul>\n<li>train: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]</a></li>\n<li>infer: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer\" target=\"_blank\">RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]</a></li>\n</ul>\n<p>Here is a simple overview of the notebook,</p>\n<h2>ROI + Rectangle Image</h2>\n<ul>\n<li>This notebook will <strong>ROI (Region Of Interest)</strong> image instead of full image. ROI is extracted using <code>OpenCV</code> (binary + max contour).</li>\n<li>This notebook will use <strong>rectangle</strong> image (width!=height; height/width=2.0) training to avoid distortion.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F0a26fb78e9c23194ae4373d784f7c9f9%2Frsna-bcd-img.png?generation=1670280891703273&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Upsample <code>Cancer</code></h2>\n<ul>\n<li>Upsample the cancer data <code>10x</code> to reduce class_imbalance effect on loss. By the way, from initial experiment it seems upsample does helps.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fd18c1a740c2f22ba376d60451d578221%2Frsna-bcd-upsample.png?generation=1670281104246304&amp;alt=media\"></li>\n</ul>\n<h2><code>TPU-1VM</code> aka <code>Local TPU</code></h2>\n<ul>\n<li>It doesn't require <code>GCS_PATH</code> like normaly <code>TPU</code>. So, for this, we don't have to make our dataset <strong>public</strong> anymore. </li>\n<li>It doesn't require <code>internet-access</code>.</li>\n<li>As datasets will be read from local path, it should speed up the first epoch (for normal <code>tpu</code> first epoch is quite slow).</li>\n</ul>\n<h2>Augmentations:</h2>\n<ul>\n<li>Random - ScaleShiftRotate</li>\n<li>Random - Horizontal Flip</li>\n<li>Random - Brightness, Contrast, Hue, Saturation</li>\n<li>Coarse Dropout</li>\n<li>MixUp - soon</li>\n</ul>\n<h2>WandB Integration:</h2>\n<ul>\n<li>You can track your training using <strong>wandb</strong></li>\n<li>It's very easy to compare model's performance using <strong>wandb</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F034e0565f957bf48e940f39a3c940e48%2Frsna-bcd-wandb.PNG?generation=1670281823925705&amp;alt=media\"></li>\n</ul>\n<h2>Grad-CAM:</h2>\n<ul>\n<li>After each fold 20 images with best and 20 images with worst results are stored on <strong>WandB</strong> for result analysis with <strong>Grad-CAM</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb3dc2945a0babb2f6239f3efe2db6f95%2Fgrad-img.png.jpg?generation=1670338916367328&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2><code>pF1</code> Maximization:</h2>\n<ul>\n<li>Use simple empirical method to attain best <code>threshold</code> for max <code>pF1</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F92572eaa1fad9dc5602fb56818b4b8dd%2Fthr.png?generation=1670338996491298&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Observations</h1>\n<ul>\n<li>Notebook require more experimentation as its result is lower than <strong>PyTorch</strong> baseline mentioned on the discussion.</li>\n<li>Models seem to overfit quietly at very early epochs hence <strong>lr-scheduler</strong> needs to adjusted</li>\n<li><strong>FocalLoss</strong> could be a viable solution to reduce class imbalance effect.</li>\n<li>Large image-size not always improves score, we need to adjust other parameters to get the better score.</li>\n<li>Upsample does improve the score but it also helps overfitting in <code>pF1</code> score.</li>\n<li>From <code>grad-cam</code> it seems model sometimes focus on the outer <strong>black</strong> region even though it doesn't contain any information.</li>\n</ul>",
  "messages": [
    {
      "id": 2056258,
      "postDate": "2022-12-05T23:27:07.290Z",
      "content": "<p>Hi everyone, there has been a new addition to <code>accelerator</code> in Kaggle <a href=\"https://www.kaggle.com/discussions/product-feedback/369338\" target=\"_blank\">here</a>, <code>TPU-1VM</code> aka <code>local TPU</code>. Unlike its predecessor <code>remote-TPU</code>,</p>\n<ul>\n<li>It doesn't require <code>GCS_PATH</code> anymore and it doesn't need <code>internet access</code>, which makes it even more easier to use. </li>\n<li>From the performance perspective unlike <code>remote-TPU</code> its first epoch is not that slow.</li>\n</ul>\n<p>I have published a notebook to show how to train in <code>TPU-1VM</code>, this notebook also supports <code>multi-GPU</code> training.</p>\n<h2>Notebooks</h2>\n<ul>\n<li>train: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train\" target=\"_blank\">RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]</a></li>\n<li>infer: <a href=\"https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer\" target=\"_blank\">RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]</a></li>\n</ul>\n<p>Here is a simple overview of the notebook,</p>\n<h2>ROI + Rectangle Image</h2>\n<ul>\n<li>This notebook will <strong>ROI (Region Of Interest)</strong> image instead of full image. ROI is extracted using <code>OpenCV</code> (binary + max contour).</li>\n<li>This notebook will use <strong>rectangle</strong> image (width!=height; height/width=2.0) training to avoid distortion.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F0a26fb78e9c23194ae4373d784f7c9f9%2Frsna-bcd-img.png?generation=1670280891703273&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2>Upsample <code>Cancer</code></h2>\n<ul>\n<li>Upsample the cancer data <code>10x</code> to reduce class_imbalance effect on loss. By the way, from initial experiment it seems upsample does helps.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fd18c1a740c2f22ba376d60451d578221%2Frsna-bcd-upsample.png?generation=1670281104246304&amp;alt=media\"></li>\n</ul>\n<h2><code>TPU-1VM</code> aka <code>Local TPU</code></h2>\n<ul>\n<li>It doesn't require <code>GCS_PATH</code> like normaly <code>TPU</code>. So, for this, we don't have to make our dataset <strong>public</strong> anymore. </li>\n<li>It doesn't require <code>internet-access</code>.</li>\n<li>As datasets will be read from local path, it should speed up the first epoch (for normal <code>tpu</code> first epoch is quite slow).</li>\n</ul>\n<h2>Augmentations:</h2>\n<ul>\n<li>Random - ScaleShiftRotate</li>\n<li>Random - Horizontal Flip</li>\n<li>Random - Brightness, Contrast, Hue, Saturation</li>\n<li>Coarse Dropout</li>\n<li>MixUp - soon</li>\n</ul>\n<h2>WandB Integration:</h2>\n<ul>\n<li>You can track your training using <strong>wandb</strong></li>\n<li>It's very easy to compare model's performance using <strong>wandb</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F034e0565f957bf48e940f39a3c940e48%2Frsna-bcd-wandb.PNG?generation=1670281823925705&amp;alt=media\"></li>\n</ul>\n<h2>Grad-CAM:</h2>\n<ul>\n<li>After each fold 20 images with best and 20 images with worst results are stored on <strong>WandB</strong> for result analysis with <strong>Grad-CAM</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb3dc2945a0babb2f6239f3efe2db6f95%2Fgrad-img.png.jpg?generation=1670338916367328&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h2><code>pF1</code> Maximization:</h2>\n<ul>\n<li>Use simple empirical method to attain best <code>threshold</code> for max <code>pF1</code><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F92572eaa1fad9dc5602fb56818b4b8dd%2Fthr.png?generation=1670338996491298&amp;alt=media\" alt=\"\"></li>\n</ul>\n<h1>Observations</h1>\n<ul>\n<li>Notebook require more experimentation as its result is lower than <strong>PyTorch</strong> baseline mentioned on the discussion.</li>\n<li>Models seem to overfit quietly at very early epochs hence <strong>lr-scheduler</strong> needs to adjusted</li>\n<li><strong>FocalLoss</strong> could be a viable solution to reduce class imbalance effect.</li>\n<li>Large image-size not always improves score, we need to adjust other parameters to get the better score.</li>\n<li>Upsample does improve the score but it also helps overfitting in <code>pF1</code> score.</li>\n<li>From <code>grad-cam</code> it seems model sometimes focus on the outer <strong>black</strong> region even though it doesn't contain any information.</li>\n</ul>",
      "rawMarkdown": "Hi everyone, there has been a new addition to `accelerator` in Kaggle [here](https://www.kaggle.com/discussions/product-feedback/369338), `TPU-1VM` aka `local TPU`. Unlike its predecessor `remote-TPU`,\n* It doesn't require `GCS_PATH` anymore and it doesn't need `internet access`, which makes it even more easier to use. \n* From the performance perspective unlike `remote-TPU` its first epoch is not that slow.\n\nI have published a notebook to show how to train in `TPU-1VM`, this notebook also supports `multi-GPU` training.\n\n## Notebooks\n* train: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n* infer: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer)\n\nHere is a simple overview of the notebook,\n\n## ROI + Rectangle Image\n* This notebook will **ROI (Region Of Interest)** image instead of full image. ROI is extracted using `OpenCV` (binary + max contour).\n* This notebook will use **rectangle** image (width!=height; height/width=2.0) training to avoid distortion.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F0a26fb78e9c23194ae4373d784f7c9f9%2Frsna-bcd-img.png?generation=1670280891703273&alt=media)\n\n## Upsample `Cancer`\n* Upsample the cancer data `10x` to reduce class_imbalance effect on loss. By the way, from initial experiment it seems upsample does helps.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fd18c1a740c2f22ba376d60451d578221%2Frsna-bcd-upsample.png?generation=1670281104246304&alt=media\" height=300>\n\n## `TPU-1VM` aka `Local TPU`\n* It doesn't require `GCS_PATH` like normaly `TPU`. So, for this, we don't have to make our dataset **public** anymore. \n* It doesn't require `internet-access`.\n* As datasets will be read from local path, it should speed up the first epoch (for normal `tpu` first epoch is quite slow).\n\n## Augmentations:\n* Random - ScaleShiftRotate\n* Random - Horizontal Flip\n* Random - Brightness, Contrast, Hue, Saturation\n* Coarse Dropout\n* MixUp - soon\n\n## WandB Integration:\n* You can track your training using **wandb**\n* It's very easy to compare model's performance using **wandb**.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F034e0565f957bf48e940f39a3c940e48%2Frsna-bcd-wandb.PNG?generation=1670281823925705&alt=media\">\n\n## Grad-CAM:\n* After each fold 20 images with best and 20 images with worst results are stored on **WandB** for result analysis with **Grad-CAM**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb3dc2945a0babb2f6239f3efe2db6f95%2Fgrad-img.png.jpg?generation=1670338916367328&alt=media)\n\n## `pF1` Maximization:\n* Use simple empirical method to attain best `threshold` for max `pF1`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F92572eaa1fad9dc5602fb56818b4b8dd%2Fthr.png?generation=1670338996491298&alt=media)\n\n# Observations\n* Notebook require more experimentation as its result is lower than **PyTorch** baseline mentioned on the discussion.\n* Models seem to overfit quietly at very early epochs hence **lr-scheduler** needs to adjusted\n* **FocalLoss** could be a viable solution to reduce class imbalance effect.\n* Large image-size not always improves score, we need to adjust other parameters to get the better score.\n* Upsample does improve the score but it also helps overfitting in `pF1` score.\n* From `grad-cam` it seems model sometimes focus on the outer **black** region even though it doesn't contain any information.",
      "votes": 34
    },
    {
      "id": 2061541,
      "postDate": "2022-12-11T08:25:04.760Z",
      "content": "<h1>Update 11 Dec 2022</h1>\n<ul>\n<li>data processing is same for both <code>train</code> and <code>test</code>, previously <code>test</code> was different due to time constraints.</li>\n<li>data fix results in <code>0.06</code> ~ <code>6%</code> improvement in <code>lb</code> (<code>0.31</code> -&gt; <code>0.37</code>)</li>\n<li>current infer notebook uses 3folds with <code>1024 x 512</code> img_size which takes <code>~7</code> hours to complete.</li>\n</ul>",
      "rawMarkdown": "# Update 11 Dec 2022\n* data processing is same for both `train` and `test`, previously `test` was different due to time constraints.\n* data fix results in `0.06` ~ `6%` improvement in `lb` (`0.31` -> `0.37`)\n* current infer notebook uses 3folds with `1024 x 512` img_size which takes `~7` hours to complete.",
      "votes": 1
    },
    {
      "id": 2059558,
      "postDate": "2022-12-09T02:39:50.247Z",
      "content": "<p>i haven't tried TPU before.<br>\ncan kaggle TPU train 2048 input model?<br>\nIf yes, up to what limit, e.g. efficient-net b2</p>",
      "rawMarkdown": "i haven't tried TPU before.\ncan kaggle TPU train 2048 input model?\nIf yes, up to what limit, e.g. efficient-net b2",
      "votes": 1,
      "replies": [
        {
          "id": 2059623,
          "postDate": "2022-12-09T04:43:08.250Z",
          "content": "<p>I guess have to experiment.</p>",
          "rawMarkdown": "I guess have to experiment."
        },
        {
          "id": 2060057,
          "postDate": "2022-12-09T14:13:52.123Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Did you train with <code>2048x2048</code> or used any <code>roi</code> method or keeping <code>aspect ratio</code> ?</p>",
          "rawMarkdown": "@hengck23 Did you train with `2048x2048` or used any `roi` method or keeping `aspect ratio` ?"
        }
      ]
    },
    {
      "id": 2058035,
      "postDate": "2022-12-07T15:05:24.670Z",
      "content": "<p>thanks for sharing, good to see a tensorflow implementation</p>",
      "rawMarkdown": "thanks for sharing, good to see a tensorflow implementation",
      "votes": 1
    },
    {
      "id": 2056265,
      "postDate": "2022-12-05T23:48:22.977Z",
      "content": "<p>Thanks for sharing this! It looks like the link to the inference notebook is missing?</p>",
      "rawMarkdown": "Thanks for sharing this! It looks like the link to the inference notebook is missing?",
      "votes": 1,
      "replies": [
        {
          "id": 2056268,
          "postDate": "2022-12-06T00:01:26.767Z",
          "content": "<p>Yes, working on it. Will publish it by today.</p>",
          "rawMarkdown": "Yes, working on it. Will publish it by today.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2056314,
      "postDate": "2022-12-06T01:44:40.937Z",
      "content": "<p>\"Notebook require more experimentation as its result is lower than PyTorch baseline mentioned on the discussion.\"</p>\n<p>check my efficient-net-b4 train logfile and other model train parameters at:    <br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug</a></p>",
      "rawMarkdown": "\"Notebook require more experimentation as its result is lower than PyTorch baseline mentioned on the discussion.\"\n\ncheck my efficient-net-b4 train logfile and other model train parameters at:    \nhttps://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug",
      "votes": 2,
      "replies": [
        {
          "id": 2056319,
          "postDate": "2022-12-06T02:04:15.743Z",
          "content": "<p>i can't tell for sure but i think your sampling (or loss) could be the problem.<br>\nYou training loss is not low enough … underfitting?</p>\n<p>my experimental results</p>\n<p><strong>note: these are initial \"improper experiments\", where both code and multiple parameters(e.g. augmentation, optimizer type) are changed during my development. Do interpret with care</strong></p>\n<p><img src=\"https://i.ibb.co/sWVT2MR/Selection-125.png\" alt=\"https://i.ibb.co/sWVT2MR/Selection-125.png\"></p>\n<p>original data pos/neg ratio = 0.02 = 1/50. you increase the pos samples by 10 times and your training data ratio is 1/5 which is too far from actual distribution. </p>\n<p>If data is perfect (not label noise and problem is easy, i.e. good separability) different distribution will not affect results, even if class is imbalance. But the kaggle ground truth here is not completely correct and the problem itself is difficult</p>",
          "rawMarkdown": "i can't tell for sure but i think your sampling (or loss) could be the problem.\nYou training loss is not low enough ... underfitting?\n\nmy experimental results\n\n**note: these are initial \"improper experiments\", where both code and multiple parameters(e.g. augmentation, optimizer type) are changed during my development. Do interpret with care**\n\n\n![https://i.ibb.co/sWVT2MR/Selection-125.png](https://i.ibb.co/sWVT2MR/Selection-125.png)\n\n\noriginal data pos/neg ratio = 0.02 = 1/50. you increase the pos samples by 10 times and your training data ratio is 1/5 which is too far from actual distribution. \n\nIf data is perfect (not label noise and problem is easy, i.e. good separability) different distribution will not affect results, even if class is imbalance. But the kaggle ground truth here is not completely correct and the problem itself is difficult",
          "votes": 1
        },
        {
          "id": 2056523,
          "postDate": "2022-12-06T07:38:50.230Z",
          "content": "<p>My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform? If model has good performance it should reflect on test regardless the imbalance. From my experiment it seems <strong>upsample</strong> does improves <code>pF1</code> and <code>auc</code>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F72ee50e51d50ac7a9bb4fdc14afe3a20%2Fupsample-result.PNG?generation=1670312301073446&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform? If model has good performance it should reflect on test regardless the imbalance. From my experiment it seems **upsample** does improves `pF1` and `auc`.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F72ee50e51d50ac7a9bb4fdc14afe3a20%2Fupsample-result.PNG?generation=1670312301073446&alt=media)"
        },
        {
          "id": 2056530,
          "postDate": "2022-12-06T07:50:00.063Z",
          "content": "<p>\"My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform?\"<br>\nit all depends on data quality. You hypothesis is correct if \"the data is clean\".</p>\n<p>Note that there is no \"sure answer\", you will have to do experiments to verify.</p>\n<hr>\n<p>as an example:</p>\n<p>assume there is \"label noise\", you can magnify or reduce this \"error\" by changing the sampling.</p>\n<p>\"label noise\" may not be be \"real noise\", it can be correct ground truth but the samples are non-separable due to the limitation of the feature space of the model </p>\n<p>the final results are combination of sampling, loss function, rate and optimizer, initialization, augmentation, etc.</p>",
          "rawMarkdown": "\"My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform?\"\nit all depends on data quality. You hypothesis is correct if \"the data is clean\".\n\nNote that there is no \"sure answer\", you will have to do experiments to verify.\n\n---\nas an example:\n\nassume there is \"label noise\", you can magnify or reduce this \"error\" by changing the sampling.\n\n\"label noise\" may not be be \"real noise\", it can be correct ground truth but the samples are non-separable due to the limitation of the feature space of the model \n\nthe final results are combination of sampling, loss function, rate and optimizer, initialization, augmentation, etc.\n\n ",
          "votes": 2
        },
        {
          "id": 2056537,
          "postDate": "2022-12-06T08:03:02.443Z",
          "content": "<p>Yes you are right, if the \"data is not clean\" it could increase the label-noise. </p>\n<p>Btw one of the reasons why I used upsample is to make sure per batch, model sees more positive samples. I guess in your code you used <code>BalanceSampler</code> which is making sure that per batch model sees positive sample. In TF I haven't used that so there is a chance that model is seeing many batches where there is no <code>pos</code> sample leading to its weights getting updated in a biased way. I wonder if that is the reason for lower score.</p>",
          "rawMarkdown": "Yes you are right, if the \"data is not clean\" it could increase the label-noise. \n\nBtw one of the reasons why I used upsample is to make sure per batch, model sees more positive samples. I guess in your code you used `BalanceSampler` which is making sure that per batch model sees positive sample. In TF I haven't used that so there is a chance that model is seeing many batches where there is no `pos` sample leading to its weights getting updated in a biased way. I wonder if that is the reason for lower score.",
          "votes": -1
        }
      ]
    },
    {
      "id": 2057547,
      "postDate": "2022-12-07T07:44:18.923Z",
      "content": "<p>This guide demonstrates how to perform basic training on Tensor Processing Units (TPUs) and TPU Pods, a collection of TPU devices connected by dedicated …New Cloud TPU VMs let you run TensorFlow, PyTorch, and JAX workloads on TPU host machines,  <a href=\"https://clickercounter.org/\" target=\"_blank\">clicker counter  </a> improving performance and usability, and reducing …This article is a detailed guide to training the popular RetinaNet object detection network on TPU. Google's Tensor Processing Units (TPUs). Finding a …This neural architecture search produced a baseline model: edgetpunet-S, which is subsequently scaled up using EfficientNet's compound scaling method to …</p>",
      "rawMarkdown": "This guide demonstrates how to perform basic training on Tensor Processing Units (TPUs) and TPU Pods, a collection of TPU devices connected by dedicated ...New Cloud TPU VMs let you run TensorFlow, PyTorch, and JAX workloads on TPU host machines,  [clicker counter  ](https://clickercounter.org/) improving performance and usability, and reducing ...This article is a detailed guide to training the popular RetinaNet object detection network on TPU. Google's Tensor Processing Units (TPUs). Finding a ...This neural architecture search produced a baseline model: edgetpunet-S, which is subsequently scaled up using EfficientNet's compound scaling method to ...",
      "votes": -3
    },
    {
      "id": 2060810,
      "postDate": "2022-12-10T13:11:20.437Z",
      "content": "<p>I didn't know that now we can use TPUs locally. Thanks for the tip</p>",
      "rawMarkdown": "I didn't know that now we can use TPUs locally. Thanks for the tip"
    },
    {
      "id": 2059612,
      "postDate": "2022-12-09T04:35:48Z",
      "content": "<h1>Update 09 Dec 2022:</h1>\n<ul>\n<li>Submission has been fixed, currently 1fold score <code>0.26</code> with thresholding.</li>\n<li>Score is likely to improve as I changed dicom processing in inference for speed up.</li>\n</ul>",
      "rawMarkdown": "# Update 09 Dec 2022:\n* Submission has been fixed, currently 1fold score `0.26` with thresholding.\n* Score is likely to improve as I changed dicom processing in inference for speed up.",
      "replies": [
        {
          "id": 2060053,
          "postDate": "2022-12-09T14:12:47.783Z",
          "content": "<ul>\n<li>2fold score is <code>0.30</code> =)</li>\n</ul>",
          "rawMarkdown": "* 2fold score is `0.30` =)"
        }
      ]
    },
    {
      "id": 2057257,
      "postDate": "2022-12-06T22:33:59.423Z",
      "content": "<p>as for Grad-CAM, check both train and test images.<br>\nif the train images are wrong, the model is learning the wrong things!</p>",
      "rawMarkdown": "as for Grad-CAM, check both train and test images.\nif the train images are wrong, the model is learning the wrong things!",
      "replies": [
        {
          "id": 2057351,
          "postDate": "2022-12-07T02:21:23.577Z",
          "content": "<p>so far, for correct predicting images grad-cams are intuitive but not sure how much correct they are clinically. Weirdly, sometimes model focuses on <strong>black-empty</strong> region, I couldn't find why…</p>",
          "rawMarkdown": "so far, for correct predicting images grad-cams are intuitive but not sure how much correct they are clinically. Weirdly, sometimes model focuses on **black-empty** region, I couldn't find why..."
        }
      ]
    },
    {
      "id": 2057240,
      "postDate": "2022-12-06T22:07:39.710Z",
      "content": "<p>your results are much improved. What is reason for the improvement?<br>\nchange of augmentation?</p>\n<p>(i note that using cropped images restrict the use of larger rotation (the rotated object will go out out of the image boundary)</p>",
      "rawMarkdown": "your results are much improved. What is reason for the improvement?\nchange of augmentation?\n\n(i note that using cropped images restrict the use of larger rotation (the rotated object will go out out of the image boundary)",
      "replies": [
        {
          "id": 2057350,
          "postDate": "2022-12-07T02:19:08.530Z",
          "content": "<p>lr increase, focal loss, vflip. But roi 5fold got timeout error last night.  Now trying with just one model. If it doesn't work, then may have to drop <strong>roi</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fdb0ba744c524fec95575770c30e2a1a3%2FCapture.PNG?generation=1670379514586813&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "lr increase, focal loss, vflip. But roi 5fold got timeout error last night.  Now trying with just one model. If it doesn't work, then may have to drop **roi**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fdb0ba744c524fec95575770c30e2a1a3%2FCapture.PNG?generation=1670379514586813&alt=media)"
        }
      ]
    },
    {
      "id": 2056569,
      "postDate": "2022-12-06T08:50:59.340Z",
      "content": "<p>Great notebook! Thank you for sharing.</p>",
      "rawMarkdown": "Great notebook! Thank you for sharing.",
      "votes": 1
    },
    {
      "id": 2056312,
      "postDate": "2022-12-06T01:44:03.490Z",
      "content": "<p>Nice, thanks for sharing.</p>",
      "rawMarkdown": "Nice, thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2060048,
      "postDate": "2022-12-09T14:11:56.940Z",
      "content": "<p>Great notebook! Thank you for sharing.</p>",
      "rawMarkdown": "Great notebook! Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 2061541,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-12-11T08:25:04.760000",
      "content": "<h1>Update 11 Dec 2022</h1>\n<ul>\n<li>data processing is same for both <code>train</code> and <code>test</code>, previously <code>test</code> was different due to time constraints.</li>\n<li>data fix results in <code>0.06</code> ~ <code>6%</code> improvement in <code>lb</code> (<code>0.31</code> -&gt; <code>0.37</code>)</li>\n<li>current infer notebook uses 3folds with <code>1024 x 512</code> img_size which takes <code>~7</code> hours to complete.</li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2059558,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-09T02:39:50.247000",
      "content": "<p>i haven't tried TPU before.<br>\ncan kaggle TPU train 2048 input model?<br>\nIf yes, up to what limit, e.g. efficient-net b2</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2059623,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-09T04:43:08.250000",
          "content": "<p>I guess have to experiment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2060057,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-09T14:13:52.123000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Did you train with <code>2048x2048</code> or used any <code>roi</code> method or keeping <code>aspect ratio</code> ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2058035,
      "author_name": "pineapple",
      "author_url": "",
      "post_date": "2022-12-07T15:05:24.670000",
      "content": "<p>thanks for sharing, good to see a tensorflow implementation</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2056265,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2022-12-05T23:48:22.977000",
      "content": "<p>Thanks for sharing this! It looks like the link to the inference notebook is missing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2056268,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-06T00:01:26.767000",
          "content": "<p>Yes, working on it. Will publish it by today.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2056314,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-06T01:44:40.937000",
      "content": "<p>\"Notebook require more experimentation as its result is lower than PyTorch baseline mentioned on the discussion.\"</p>\n<p>check my efficient-net-b4 train logfile and other model train parameters at:    <br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2056319,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-06T02:04:15.743000",
          "content": "<p>i can't tell for sure but i think your sampling (or loss) could be the problem.<br>\nYou training loss is not low enough … underfitting?</p>\n<p>my experimental results</p>\n<p><strong>note: these are initial \"improper experiments\", where both code and multiple parameters(e.g. augmentation, optimizer type) are changed during my development. Do interpret with care</strong></p>\n<p><img src=\"https://i.ibb.co/sWVT2MR/Selection-125.png\" alt=\"https://i.ibb.co/sWVT2MR/Selection-125.png\"></p>\n<p>original data pos/neg ratio = 0.02 = 1/50. you increase the pos samples by 10 times and your training data ratio is 1/5 which is too far from actual distribution. </p>\n<p>If data is perfect (not label noise and problem is easy, i.e. good separability) different distribution will not affect results, even if class is imbalance. But the kaggle ground truth here is not completely correct and the problem itself is difficult</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2056523,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-06T07:38:50.230000",
          "content": "<p>My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform? If model has good performance it should reflect on test regardless the imbalance. From my experiment it seems <strong>upsample</strong> does improves <code>pF1</code> and <code>auc</code>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F72ee50e51d50ac7a9bb4fdc14afe3a20%2Fupsample-result.PNG?generation=1670312301073446&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2056530,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-12-06T07:50:00.063000",
          "content": "<p>\"My hypothesis was if a model is trained 50/50 dataset then tested on 90/10 dataset how will it perform?\"<br>\nit all depends on data quality. You hypothesis is correct if \"the data is clean\".</p>\n<p>Note that there is no \"sure answer\", you will have to do experiments to verify.</p>\n<hr>\n<p>as an example:</p>\n<p>assume there is \"label noise\", you can magnify or reduce this \"error\" by changing the sampling.</p>\n<p>\"label noise\" may not be be \"real noise\", it can be correct ground truth but the samples are non-separable due to the limitation of the feature space of the model </p>\n<p>the final results are combination of sampling, loss function, rate and optimizer, initialization, augmentation, etc.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2056537,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-06T08:03:02.443000",
          "content": "<p>Yes you are right, if the \"data is not clean\" it could increase the label-noise. </p>\n<p>Btw one of the reasons why I used upsample is to make sure per batch, model sees more positive samples. I guess in your code you used <code>BalanceSampler</code> which is making sure that per batch model sees positive sample. In TF I haven't used that so there is a chance that model is seeing many batches where there is no <code>pos</code> sample leading to its weights getting updated in a biased way. I wonder if that is the reason for lower score.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 2057547,
      "author_name": "Nora Davis",
      "author_url": "",
      "post_date": "2022-12-07T07:44:18.923000",
      "content": "<p>This guide demonstrates how to perform basic training on Tensor Processing Units (TPUs) and TPU Pods, a collection of TPU devices connected by dedicated …New Cloud TPU VMs let you run TensorFlow, PyTorch, and JAX workloads on TPU host machines,  <a href=\"https://clickercounter.org/\" target=\"_blank\">clicker counter  </a> improving performance and usability, and reducing …This article is a detailed guide to training the popular RetinaNet object detection network on TPU. Google's Tensor Processing Units (TPUs). Finding a …This neural architecture search produced a baseline model: edgetpunet-S, which is subsequently scaled up using EfficientNet's compound scaling method to …</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 2060810,
      "author_name": "Robson",
      "author_url": "",
      "post_date": "2022-12-10T13:11:20.437000",
      "content": "<p>I didn't know that now we can use TPUs locally. Thanks for the tip</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2059612,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-12-09T04:35:48",
      "content": "<h1>Update 09 Dec 2022:</h1>\n<ul>\n<li>Submission has been fixed, currently 1fold score <code>0.26</code> with thresholding.</li>\n<li>Score is likely to improve as I changed dicom processing in inference for speed up.</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2060053,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-09T14:12:47.783000",
          "content": "<ul>\n<li>2fold score is <code>0.30</code> =)</li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2057257,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-06T22:33:59.423000",
      "content": "<p>as for Grad-CAM, check both train and test images.<br>\nif the train images are wrong, the model is learning the wrong things!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2057351,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-07T02:21:23.577000",
          "content": "<p>so far, for correct predicting images grad-cams are intuitive but not sure how much correct they are clinically. Weirdly, sometimes model focuses on <strong>black-empty</strong> region, I couldn't find why…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2057240,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-06T22:07:39.710000",
      "content": "<p>your results are much improved. What is reason for the improvement?<br>\nchange of augmentation?</p>\n<p>(i note that using cropped images restrict the use of larger rotation (the rotated object will go out out of the image boundary)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2057350,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-12-07T02:19:08.530000",
          "content": "<p>lr increase, focal loss, vflip. But roi 5fold got timeout error last night.  Now trying with just one model. If it doesn't work, then may have to drop <strong>roi</strong>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fdb0ba744c524fec95575770c30e2a1a3%2FCapture.PNG?generation=1670379514586813&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2056569,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-06T08:50:59.340000",
      "content": "<p>Great notebook! Thank you for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2056312,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2022-12-06T01:44:03.490000",
      "content": "<p>Nice, thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2060048,
      "author_name": "Pavithra Devi M",
      "author_url": "",
      "post_date": "2022-12-09T14:11:56.940000",
      "content": "<p>Great notebook! Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2056258": "Hi everyone, there has been a new addition to `accelerator` in Kaggle [here](https://www.kaggle.com/discussions/product-feedback/369338), `TPU-1VM` aka `local TPU`. Unlike its predecessor `remote-TPU`,\n* It doesn't require `GCS_PATH` anymore and it doesn't need `internet access`, which makes it even more easier to use. \n* From the performance perspective unlike `remote-TPU` its first epoch is not that slow.\n\nI have published a notebook to show how to train in `TPU-1VM`, this notebook also supports `multi-GPU` training.\n\n## Notebooks\n* train: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Train]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-train)\n* infer: [RSNA-BCD: EfficientNet [TF][TPU-1VM][Infer]](https://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer)\n\nHere is a simple overview of the notebook,\n\n## ROI + Rectangle Image\n* This notebook will **ROI (Region Of Interest)** image instead of full image. ROI is extracted using `OpenCV` (binary + max contour).\n* This notebook will use **rectangle** image (width!=height; height/width=2.0) training to avoid distortion.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F0a26fb78e9c23194ae4373d784f7c9f9%2Frsna-bcd-img.png?generation=1670280891703273&alt=media)\n\n## Upsample `Cancer`\n* Upsample the cancer data `10x` to reduce class_imbalance effect on loss. By the way, from initial experiment it seems upsample does helps.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fd18c1a740c2f22ba376d60451d578221%2Frsna-bcd-upsample.png?generation=1670281104246304&alt=media\" height=300>\n\n## `TPU-1VM` aka `Local TPU`\n* It doesn't require `GCS_PATH` like normaly `TPU`. So, for this, we don't have to make our dataset **public** anymore. \n* It doesn't require `internet-access`.\n* As datasets will be read from local path, it should speed up the first epoch (for normal `tpu` first epoch is quite slow).\n\n## Augmentations:\n* Random - ScaleShiftRotate\n* Random - Horizontal Flip\n* Random - Brightness, Contrast, Hue, Saturation\n* Coarse Dropout\n* MixUp - soon\n\n## WandB Integration:\n* You can track your training using **wandb**\n* It's very easy to compare model's performance using **wandb**.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F034e0565f957bf48e940f39a3c940e48%2Frsna-bcd-wandb.PNG?generation=1670281823925705&alt=media\">\n\n## Grad-CAM:\n* After each fold 20 images with best and 20 images with worst results are stored on **WandB** for result analysis with **Grad-CAM**.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2Fb3dc2945a0babb2f6239f3efe2db6f95%2Fgrad-img.png.jpg?generation=1670338916367328&alt=media)\n\n## `pF1` Maximization:\n* Use simple empirical method to attain best `threshold` for max `pF1`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3574256%2F92572eaa1fad9dc5602fb56818b4b8dd%2Fthr.png?generation=1670338996491298&alt=media)\n\n# Observations\n* Notebook require more experimentation as its result is lower than **PyTorch** baseline mentioned on the discussion.\n* Models seem to overfit quietly at very early epochs hence **lr-scheduler** needs to adjusted\n* **FocalLoss** could be a viable solution to reduce class imbalance effect.\n* Large image-size not always improves score, we need to adjust other parameters to get the better score.\n* Upsample does improve the score but it also helps overfitting in `pF1` score.\n* From `grad-cam` it seems model sometimes focus on the outer **black** region even though it doesn't contain any information.",
    "2061541": "# Update 11 Dec 2022\n* data processing is same for both `train` and `test`, previously `test` was different due to time constraints.\n* data fix results in `0.06` ~ `6%` improvement in `lb` (`0.31` -> `0.37`)\n* current infer notebook uses 3folds with `1024 x 512` img_size which takes `~7` hours to complete.",
    "2059558": "i haven't tried TPU before.\ncan kaggle TPU train 2048 input model?\nIf yes, up to what limit, e.g. efficient-net b2",
    "2058035": "thanks for sharing, good to see a tensorflow implementation",
    "2056265": "Thanks for sharing this! It looks like the link to the inference notebook is missing?",
    "2056314": "\"Notebook require more experimentation as its result is lower than PyTorch baseline mentioned on the discussion.\"\n\ncheck my efficient-net-b4 train logfile and other model train parameters at:    \nhttps://www.kaggle.com/datasets/hengck23/for-tpu-efficientb4-debug",
    "2057547": "This guide demonstrates how to perform basic training on Tensor Processing Units (TPUs) and TPU Pods, a collection of TPU devices connected by dedicated ...New Cloud TPU VMs let you run TensorFlow, PyTorch, and JAX workloads on TPU host machines,  [clicker counter  ](https://clickercounter.org/) improving performance and usability, and reducing ...This article is a detailed guide to training the popular RetinaNet object detection network on TPU. Google's Tensor Processing Units (TPUs). Finding a ...This neural architecture search produced a baseline model: edgetpunet-S, which is subsequently scaled up using EfficientNet's compound scaling method to ...",
    "2060810": "I didn't know that now we can use TPUs locally. Thanks for the tip",
    "2059612": "# Update 09 Dec 2022:\n* Submission has been fixed, currently 1fold score `0.26` with thresholding.\n* Score is likely to improve as I changed dicom processing in inference for speed up.",
    "2057257": "as for Grad-CAM, check both train and test images.\nif the train images are wrong, the model is learning the wrong things!",
    "2057240": "your results are much improved. What is reason for the improvement?\nchange of augmentation?\n\n(i note that using cropped images restrict the use of larger rotation (the rotated object will go out out of the image boundary)",
    "2056569": "Great notebook! Thank you for sharing.",
    "2056312": "Nice, thanks for sharing.",
    "2060048": "Great notebook! Thank you for sharing."
  }
}