{
  "id": 219925,
  "title": "YOLOv5 Training Issue",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/219925",
  "author_name": "Daniel Hagan",
  "post_date": "2021-02-17T00:09:41.112000",
  "votes": 6,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Currently we are having issues training a YOLOv5 (using Adam optimizer) network on the data. Even with shuffling, it will only train up to step 26 (if you use batch size = 180 otherwise it will be step 51 with batch size = 90) where it unexpectedly stops, no error message is given as if the program ran successfully… Is anyone else having a similar issue? Here is a link to the code we started with<a href=\"url\" target=\"_blank\"> https://github.com/jahongir7174/YOLOv5-tf</a></p>",
  "messages": [
    {
      "id": 1205738,
      "postDate": "2021-02-17T00:09:41.113Z",
      "content": "<p>Currently we are having issues training a YOLOv5 (using Adam optimizer) network on the data. Even with shuffling, it will only train up to step 26 (if you use batch size = 180 otherwise it will be step 51 with batch size = 90) where it unexpectedly stops, no error message is given as if the program ran successfully… Is anyone else having a similar issue? Here is a link to the code we started with<a href=\"url\" target=\"_blank\"> https://github.com/jahongir7174/YOLOv5-tf</a></p>",
      "rawMarkdown": "Currently we are having issues training a YOLOv5 (using Adam optimizer) network on the data. Even with shuffling, it will only train up to step 26 (if you use batch size = 180 otherwise it will be step 51 with batch size = 90) where it unexpectedly stops, no error message is given as if the program ran successfully... Is anyone else having a similar issue? Here is a link to the code we started with[ https://github.com/jahongir7174/YOLOv5-tf](url)",
      "votes": 6
    },
    {
      "id": 1210789,
      "postDate": "2021-02-19T17:25:05.390Z",
      "content": "<p>Did you install and configure wandb properly? </p>\n<p>We trained YOLOv5 on Colab without problem.<br>\n==training==<br>\n<code>!python train.py --img 512 --batch 16 --epochs 60 --data chest-abnormal/vinbigdata.yaml --weights yolov5x.pt --cache</code><br>\n==vinbigdata.yaml==,</p>\n<p>yaml:<br>\nnames:</p>\n<ul>\n<li>Aortic enlargement</li>\n<li>Atelectasis</li>\n<li>Calcification</li>\n<li>Cardiomegaly</li>\n<li>Consolidation</li>\n<li>ILD</li>\n<li>Infiltration</li>\n<li>Lung Opacity</li>\n<li>Nodule/Mass</li>\n<li>Other lesion</li>\n<li>Pleural effusion</li>\n<li>Pleural thickening</li>\n<li>Pneumothorax</li>\n<li>Pulmonary fibrosis<br>\nnc: 14<br>\ntrain: chest-abnormal/train.txt<br>\nval: chest-abnormal/val.txt</li>\n</ul>\n<p>==log==<br>\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size<br>\n     57/59     9.82G   0.03037    0.0175   0.01142   0.05929        25       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.54it/s]<br>\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10&lt;00:00,  5.11it/s]<br>\n                 all         879    7.22e+03       0.403       0.273       0.232       0.101</p>\n<pre><code> Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n 58/59     9.82G   0.03001   0.01708   0.01122   0.05831        36       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.53it/s]\n           Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10&lt;00:00,  5.14it/s]\n             all         879    7.22e+03       0.393       0.275        0.23       0.103\n\n Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n 59/59     9.82G   0.02999   0.01684   0.01104   0.05787        52       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.53it/s]\n           Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:13&lt;00:00,  3.97it/s]\n</code></pre>",
      "rawMarkdown": "Did you install and configure wandb properly? \n\nWe trained YOLOv5 on Colab without problem.\n==training==\n`!python train.py --img 512 --batch 16 --epochs 60 --data chest-abnormal/vinbigdata.yaml --weights yolov5x.pt --cache `\n==vinbigdata.yaml==,\n\nyaml:\nnames:\n- Aortic enlargement\n- Atelectasis\n- Calcification\n- Cardiomegaly\n- Consolidation\n- ILD\n- Infiltration\n- Lung Opacity\n- Nodule/Mass\n- Other lesion\n- Pleural effusion\n- Pleural thickening\n- Pneumothorax\n- Pulmonary fibrosis\nnc: 14\ntrain: chest-abnormal/train.txt\nval: chest-abnormal/val.txt\n\n\n==log==\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     57/59     9.82G   0.03037    0.0175   0.01142   0.05929        25       512: 100%|██████████| 220/220 [01:02<00:00,  3.54it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10<00:00,  5.11it/s]\n                 all         879    7.22e+03       0.403       0.273       0.232       0.101\n\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     58/59     9.82G   0.03001   0.01708   0.01122   0.05831        36       512: 100%|██████████| 220/220 [01:02<00:00,  3.53it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10<00:00,  5.14it/s]\n                 all         879    7.22e+03       0.393       0.275        0.23       0.103\n\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     59/59     9.82G   0.02999   0.01684   0.01104   0.05787        52       512: 100%|██████████| 220/220 [01:02<00:00,  3.53it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:13<00:00,  3.97it/s]\n",
      "votes": 3,
      "replies": [
        {
          "id": 1212176,
          "postDate": "2021-02-21T00:56:35.570Z",
          "content": "<p>Thanks for the input, we actually solved the problem: it had to do with batch sizes and a cache that was being used, once it had gone through enough batches, the last batch would be a little smaller than the others and it wasn't fitting in the cache</p>",
          "rawMarkdown": "Thanks for the input, we actually solved the problem: it had to do with batch sizes and a cache that was being used, once it had gone through enough batches, the last batch would be a little smaller than the others and it wasn't fitting in the cache"
        },
        {
          "id": 1214271,
          "postDate": "2021-02-22T17:39:43.337Z",
          "content": "<p>interesting, thanks for the updates.</p>",
          "rawMarkdown": "interesting, thanks for the updates.",
          "votes": 1
        },
        {
          "id": 1217316,
          "postDate": "2021-02-25T00:01:11.837Z",
          "content": "<p>Amazing! Is  Only 512 size  so good that its lb score ranked 0.270 ? </p>",
          "rawMarkdown": "Amazing! Is  Only 512 size  so good that its lb score ranked 0.270 ? "
        },
        {
          "id": 1220224,
          "postDate": "2021-02-27T18:35:30.863Z",
          "content": "<p>single model with 1024 can get me to about LB 0.18<br>\nwith 2 class filter to about LB 0.23<br>\nensemble models to about LB 0.27</p>",
          "rawMarkdown": "single model with 1024 can get me to about LB 0.18\nwith 2 class filter to about LB 0.23\nensemble models to about LB 0.27",
          "votes": 1
        },
        {
          "id": 1220409,
          "postDate": "2021-02-28T00:17:05.067Z",
          "content": "<p>Thank you ! It is a good strategy without heavy cost.</p>",
          "rawMarkdown": "Thank you ! It is a good strategy without heavy cost."
        },
        {
          "id": 1222702,
          "postDate": "2021-03-02T04:18:52.107Z",
          "content": "<p><a href=\"https://www.kaggle.com/dennywangdev\" target=\"_blank\">@dennywangdev</a> what conf thresholds and yolo model are you using? I have found a threshold of 0 works well and have tried yolo5x. </p>",
          "rawMarkdown": "@dennywangdev what conf thresholds and yolo model are you using? I have found a threshold of 0 works well and have tried yolo5x. "
        },
        {
          "id": 1232805,
          "postDate": "2021-03-10T02:10:33.697Z",
          "content": "<p>yolov5x6.pt, conf=0.06 <br>\nanother model is detectron x101, threshold = 0.0</p>",
          "rawMarkdown": "yolov5x6.pt, conf=0.06 \nanother model is detectron x101, threshold = 0.0",
          "votes": 1
        },
        {
          "id": 1233051,
          "postDate": "2021-03-10T05:30:17.783Z",
          "content": "<p>Hello,when you ensemble, do you make 2 class filter each model before ensemble  or ensemble them before 2 class filter  ?</p>",
          "rawMarkdown": "Hello,when you ensemble, do you make 2 class filter each model before ensemble  or ensemble them before 2 class filter  ?"
        },
        {
          "id": 1233083,
          "postDate": "2021-03-10T06:02:20.987Z",
          "content": "<p>It makes no difference if you are using the same filter ( at least to LB score).</p>",
          "rawMarkdown": "It makes no difference if you are using the same filter ( at least to LB score).",
          "votes": 1
        },
        {
          "id": 1233130,
          "postDate": "2021-03-10T06:58:52.617Z",
          "content": "<p>Thanks,I think use filter in the end will be fine.</p>",
          "rawMarkdown": "Thanks,I think use filter in the end will be fine."
        },
        {
          "id": 1233727,
          "postDate": "2021-03-10T16:31:56.707Z",
          "content": "<p>Hi, Denny, I use yolov3 (with 2 class filter) and have a LB score 2.3. I am trying with yolo5X now, but only got LB around 2.0. I guess yolo5x should achieve better performance than yolov3. Do you know any possible reasons? FYI, I already enable the \"multi-label\" for nms in yolov5. Thanks!    </p>",
          "rawMarkdown": "Hi, Denny, I use yolov3 (with 2 class filter) and have a LB score 2.3. I am trying with yolo5X now, but only got LB around 2.0. I guess yolo5x should achieve better performance than yolov3. Do you know any possible reasons? FYI, I already enable the \"multi-label\" for nms in yolov5. Thanks!    "
        },
        {
          "id": 1233744,
          "postDate": "2021-03-10T16:36:36.297Z",
          "content": "<p>One more question, I'd like to try your strategy. How is \"ensemble models\" performed?</p>",
          "rawMarkdown": "One more question, I'd like to try your strategy. How is \"ensemble models\" performed?"
        },
        {
          "id": 1234302,
          "postDate": "2021-03-11T05:48:14.987Z",
          "content": "<p>From my testing, yolo5x is better at large image, did you resize your image to 1024? Plus I use yolo5x6, a variation of yolo5x. you can find weight file <a href=\"https://github.com/ultralytics/yolov5/releases\" target=\"_blank\">here</a> .</p>\n<p>Tried a few ensemble methods, the one works so far is removing the exact same bboxes(same class + IOU &gt; 0.95) then adding the rest bboxes together while keeps conf score.</p>",
          "rawMarkdown": "From my testing, yolo5x is better at large image, did you resize your image to 1024? Plus I use yolo5x6, a variation of yolo5x. you can find weight file [here](https://github.com/ultralytics/yolov5/releases) .\n\nTried a few ensemble methods, the one works so far is removing the exact same bboxes(same class + IOU > 0.95) then adding the rest bboxes together while keeps conf score.",
          "votes": 2
        },
        {
          "id": 1234661,
          "postDate": "2021-03-11T13:18:47.473Z",
          "content": "<p>Cool, thanks! Will try your ensemble strategy. Yeah, I train yolo5x on 1024, and use the default augmentations. Are you using any EXTRA augmentations (to enlarge the dataset for yolov5) or label fusion (from 3 radiologists) for preprocessing? </p>",
          "rawMarkdown": "Cool, thanks! Will try your ensemble strategy. Yeah, I train yolo5x on 1024, and use the default augmentations. Are you using any EXTRA augmentations (to enlarge the dataset for yolov5) or label fusion (from 3 radiologists) for preprocessing? "
        },
        {
          "id": 1237315,
          "postDate": "2021-03-14T03:39:24.453Z",
          "content": "<p>Good luck. </p>\n<p>And about augmentation, I tried from very little to very heavy augmentations, heavy augmentations improved CV not LB, I have to turn down aug to improve LB score a little bit. I'm confused. </p>",
          "rawMarkdown": "Good luck. \n\nAnd about augmentation, I tried from very little to very heavy augmentations, heavy augmentations improved CV not LB, I have to turn down aug to improve LB score a little bit. I'm confused. ",
          "votes": 1
        },
        {
          "id": 1238069,
          "postDate": "2021-03-14T16:05:32.893Z",
          "content": "<p>Thanks for all the feedback Denny, it has been helpful…. but for some reason our model performs quite terribly (LB Score 0.05x), which is hardly better than submitting a file of all no findings…. we check the boxes from the submission file back to the original images and they look reasonable… On the test data we looked at the predicted boxes literally fall right on top of most radiologists findings. We also got rid of overlapping boxes of the same class IOU &gt; 0.9 and somehow got the same score as before we removed them… We found a program to calculate mAP on the training set so we can at least see how it would score us there…. This is so frustrating! How can we be using the same model. I understand how to ensemble the models and do folds etc. but no reason to try any of that with the LB score we currently have. Will let everyone know what our mAP calculation on the training data comes out to… it should be high looking at the boxes, I really am not sure what is going on/to try next if anyone has suggestions…</p>",
          "rawMarkdown": "Thanks for all the feedback Denny, it has been helpful.... but for some reason our model performs quite terribly (LB Score 0.05x), which is hardly better than submitting a file of all no findings.... we check the boxes from the submission file back to the original images and they look reasonable... On the test data we looked at the predicted boxes literally fall right on top of most radiologists findings. We also got rid of overlapping boxes of the same class IOU > 0.9 and somehow got the same score as before we removed them... We found a program to calculate mAP on the training set so we can at least see how it would score us there.... This is so frustrating! How can we be using the same model. I understand how to ensemble the models and do folds etc. but no reason to try any of that with the LB score we currently have. Will let everyone know what our mAP calculation on the training data comes out to... it should be high looking at the boxes, I really am not sure what is going on/to try next if anyone has suggestions..."
        },
        {
          "id": 1239624,
          "postDate": "2021-03-15T21:18:07.520Z",
          "content": "<p>It's hard to tell what's missing in your case without seeing your notebook.</p>\n<p>For my experience, after 15 epochs, normally mAP:50 = ~0.4, mAP:0.5:0.95 = ~0.2, and LB = ~0.18 (without 2 classes filter)</p>\n<p>And if you specify val data in the data yaml,  Yolo script should be able to generate mAP:50 and mAP:0.5:0.95 and other metrics automatically. </p>\n<p>my yaml seen below</p>\n<p>==data.yaml==</p>\n<p>names:</p>\n<ul>\n<li>Aortic enlargement</li>\n<li>Atelectasis</li>\n<li>Calcification</li>\n<li>Cardiomegaly</li>\n<li>Consolidation</li>\n<li>ILD</li>\n<li>Infiltration</li>\n<li>Lung Opacity</li>\n<li>Nodule/Mass</li>\n<li>Other lesion</li>\n<li>Pleural effusion</li>\n<li>Pleural thickening</li>\n<li>Pneumothorax</li>\n<li>Pulmonary fibrosis<br>\nnc: 14<br>\ntrain: chest-abnormal/train.txt<br>\nval: chest-abnormal/val.txt</li>\n</ul>\n<p>==end==</p>\n<p>==val.txt==<br>\nchest-abnormal/yolo/images/val/78057413f8f7e8a3b1decc815e1f509e.png<br>\nchest-abnormal/yolo/images/val/d625684a437d0b9f622fec5329c6d7af.png<br>\nchest-abnormal/yolo/images/val/163898fbc57f00f58ad27e72031a541f.png<br>\nchest-abnormal/yolo/images/val/d9245845860a540560404f47bcda1716.png<br>\nchest-abnormal/yolo/images/val/7729fbb58b8006fe2b8305f9b2b33883.png<br>\nchest-abnormal/yolo/images/val/95b32ce2a10f57629eb63830376237ca.png<br>\nchest-abnormal/yolo/images/val/723f93a3f0e3905c498839d23b961631.png<br>\n…<br>\n==end==</p>",
          "rawMarkdown": "It's hard to tell what's missing in your case without seeing your notebook.\n\nFor my experience, after 15 epochs, normally mAP:50 = ~0.4, mAP:0.5:0.95 = ~0.2, and LB = ~0.18 (without 2 classes filter)\n\nAnd if you specify val data in the data yaml,  Yolo script should be able to generate mAP:50 and mAP:0.5:0.95 and other metrics automatically. \n\nmy yaml seen below\n\n==data.yaml==\n\nnames:\n- Aortic enlargement\n- Atelectasis\n- Calcification\n- Cardiomegaly\n- Consolidation\n- ILD\n- Infiltration\n- Lung Opacity\n- Nodule/Mass\n- Other lesion\n- Pleural effusion\n- Pleural thickening\n- Pneumothorax\n- Pulmonary fibrosis\nnc: 14\ntrain: chest-abnormal/train.txt\nval: chest-abnormal/val.txt\n\n==end==\n\n==val.txt==\nchest-abnormal/yolo/images/val/78057413f8f7e8a3b1decc815e1f509e.png\nchest-abnormal/yolo/images/val/d625684a437d0b9f622fec5329c6d7af.png\nchest-abnormal/yolo/images/val/163898fbc57f00f58ad27e72031a541f.png\nchest-abnormal/yolo/images/val/d9245845860a540560404f47bcda1716.png\nchest-abnormal/yolo/images/val/7729fbb58b8006fe2b8305f9b2b33883.png\nchest-abnormal/yolo/images/val/95b32ce2a10f57629eb63830376237ca.png\nchest-abnormal/yolo/images/val/723f93a3f0e3905c498839d23b961631.png\n...\n==end=="
        }
      ]
    },
    {
      "id": 1215896,
      "postDate": "2021-02-24T03:05:12.657Z",
      "content": "<p>Hello,how much score can yolov5 get?</p>",
      "rawMarkdown": "Hello,how much score can yolov5 get?",
      "votes": -1,
      "replies": [
        {
          "id": 1217226,
          "postDate": "2021-02-24T21:18:45.853Z",
          "content": "<p>I have not been able to make a submission file even though I have model trained, just getting the output in the right format so I can get back to you when we get a baseline score, but there are lots of people using this network, and it is one of the most successful objection detection models around these days. I would look at the leaderboard, if you noticed, Denny Wang who commented above is also using yolov5 and currently 16th on the LB… which is a score of 0.272</p>",
          "rawMarkdown": "I have not been able to make a submission file even though I have model trained, just getting the output in the right format so I can get back to you when we get a baseline score, but there are lots of people using this network, and it is one of the most successful objection detection models around these days. I would look at the leaderboard, if you noticed, Denny Wang who commented above is also using yolov5 and currently 16th on the LB... which is a score of 0.272"
        },
        {
          "id": 1217313,
          "postDate": "2021-02-24T23:59:10.137Z",
          "content": "<p>Thank you, it is an efficient method.I do not have enough gpus.</p>",
          "rawMarkdown": "Thank you, it is an efficient method.I do not have enough gpus."
        },
        {
          "id": 1217376,
          "postDate": "2021-02-25T01:59:03.283Z",
          "content": "<p>Right, one can never have too many GPUs xD</p>",
          "rawMarkdown": "Right, one can never have too many GPUs xD",
          "votes": -1
        },
        {
          "id": 1217384,
          "postDate": "2021-02-25T02:08:42.767Z",
          "content": "<p>We can run it on Kaggle,single fold scored 0.15x on lb :)</p>",
          "rawMarkdown": "We can run it on Kaggle,single fold scored 0.15x on lb :)"
        },
        {
          "id": 1218188,
          "postDate": "2021-02-25T16:24:31.423Z",
          "content": "<p>Awesome, not too bad! We fixed our problem compiling the model so we should have a submission in of our own today too!</p>",
          "rawMarkdown": "Awesome, not too bad! We fixed our problem compiling the model so we should have a submission in of our own today too!"
        }
      ]
    },
    {
      "id": 1205750,
      "postDate": "2021-02-17T00:18:56.713Z",
      "content": "<p><strong>ISSUE RESOLVED</strong></p>\n<p>Here is the output from PyCharm we print out the loss and the gradient (one from each GPU):<br>\n23/73 [========&gt;…………………] - ETA: 4:54 - loss: 24.2436loss =tf.Tensor(24.243551, shape=(), dtype=float32)<br>\ngrad: 44.53920408875909<br>\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)<br>\ngrad: 44.539202217558866<br>\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)<br>\n24/73 [========&gt;…………………] - ETA: 4:48 - loss: 23.7136loss =tf.Tensor(23.713629, shape=(), dtype=float32)<br>\ngrad: 61.30926377767098<br>\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)<br>\ngrad: 61.30925365406151<br>\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)<br>\n25/73 [=========&gt;………………..] - ETA: 4:40 - loss: 23.2499loss =tf.Tensor(23.249947, shape=(), dtype=float32)<br>\ngrad: 1.3752864851022184<br>\nloss: tf.Tensor(1.1149328, shape=(), dtype=float32)<br>\ngrad: 56.96218364695795<br>\nloss: tf.Tensor(0.97858346, shape=(), dtype=float32)<br>\n26/73 [=========&gt;………………..] - ETA: 4:28 - loss: 2.0935 loss =tf.Tensor(2.0935163, shape=(), dtype=float32)</p>\n<p>Process finished with exit code 0</p>",
      "rawMarkdown": "**ISSUE RESOLVED**\n\nHere is the output from PyCharm we print out the loss and the gradient (one from each GPU):\n23/73 [========>.....................] - ETA: 4:54 - loss: 24.2436loss =tf.Tensor(24.243551, shape=(), dtype=float32)\ngrad: 44.53920408875909\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)\ngrad: 44.539202217558866\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)\n24/73 [========>.....................] - ETA: 4:48 - loss: 23.7136loss =tf.Tensor(23.713629, shape=(), dtype=float32)\ngrad: 61.30926377767098\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)\ngrad: 61.30925365406151\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)\n25/73 [=========>....................] - ETA: 4:40 - loss: 23.2499loss =tf.Tensor(23.249947, shape=(), dtype=float32)\ngrad: 1.3752864851022184\nloss: tf.Tensor(1.1149328, shape=(), dtype=float32)\ngrad: 56.96218364695795\nloss: tf.Tensor(0.97858346, shape=(), dtype=float32)\n26/73 [=========>....................] - ETA: 4:28 - loss: 2.0935 loss =tf.Tensor(2.0935163, shape=(), dtype=float32)\n\nProcess finished with exit code 0"
    }
  ],
  "comments": [
    {
      "id": 1210789,
      "author_name": "Denny Wang",
      "author_url": "",
      "post_date": "2021-02-19T17:25:05.390000",
      "content": "<p>Did you install and configure wandb properly? </p>\n<p>We trained YOLOv5 on Colab without problem.<br>\n==training==<br>\n<code>!python train.py --img 512 --batch 16 --epochs 60 --data chest-abnormal/vinbigdata.yaml --weights yolov5x.pt --cache</code><br>\n==vinbigdata.yaml==,</p>\n<p>yaml:<br>\nnames:</p>\n<ul>\n<li>Aortic enlargement</li>\n<li>Atelectasis</li>\n<li>Calcification</li>\n<li>Cardiomegaly</li>\n<li>Consolidation</li>\n<li>ILD</li>\n<li>Infiltration</li>\n<li>Lung Opacity</li>\n<li>Nodule/Mass</li>\n<li>Other lesion</li>\n<li>Pleural effusion</li>\n<li>Pleural thickening</li>\n<li>Pneumothorax</li>\n<li>Pulmonary fibrosis<br>\nnc: 14<br>\ntrain: chest-abnormal/train.txt<br>\nval: chest-abnormal/val.txt</li>\n</ul>\n<p>==log==<br>\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size<br>\n     57/59     9.82G   0.03037    0.0175   0.01142   0.05929        25       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.54it/s]<br>\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10&lt;00:00,  5.11it/s]<br>\n                 all         879    7.22e+03       0.403       0.273       0.232       0.101</p>\n<pre><code> Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n 58/59     9.82G   0.03001   0.01708   0.01122   0.05831        36       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.53it/s]\n           Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10&lt;00:00,  5.14it/s]\n             all         879    7.22e+03       0.393       0.275        0.23       0.103\n\n Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n 59/59     9.82G   0.02999   0.01684   0.01104   0.05787        52       512: 100%|██████████| 220/220 [01:02&lt;00:00,  3.53it/s]\n           Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:13&lt;00:00,  3.97it/s]\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1212176,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-02-21T00:56:35.570000",
          "content": "<p>Thanks for the input, we actually solved the problem: it had to do with batch sizes and a cache that was being used, once it had gone through enough batches, the last batch would be a little smaller than the others and it wasn't fitting in the cache</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1214271,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-02-22T17:39:43.337000",
          "content": "<p>interesting, thanks for the updates.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1217316,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-25T00:01:11.837000",
          "content": "<p>Amazing! Is  Only 512 size  so good that its lb score ranked 0.270 ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220224,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-02-27T18:35:30.863000",
          "content": "<p>single model with 1024 can get me to about LB 0.18<br>\nwith 2 class filter to about LB 0.23<br>\nensemble models to about LB 0.27</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220409,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-28T00:17:05.067000",
          "content": "<p>Thank you ! It is a good strategy without heavy cost.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222702,
          "author_name": "Trushant Kalyanpur",
          "author_url": "",
          "post_date": "2021-03-02T04:18:52.107000",
          "content": "<p><a href=\"https://www.kaggle.com/dennywangdev\" target=\"_blank\">@dennywangdev</a> what conf thresholds and yolo model are you using? I have found a threshold of 0 works well and have tried yolo5x. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1232805,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-03-10T02:10:33.697000",
          "content": "<p>yolov5x6.pt, conf=0.06 <br>\nanother model is detectron x101, threshold = 0.0</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233051,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-03-10T05:30:17.783000",
          "content": "<p>Hello,when you ensemble, do you make 2 class filter each model before ensemble  or ensemble them before 2 class filter  ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233083,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-03-10T06:02:20.987000",
          "content": "<p>It makes no difference if you are using the same filter ( at least to LB score).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1233130,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-03-10T06:58:52.617000",
          "content": "<p>Thanks,I think use filter in the end will be fine.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233727,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-10T16:31:56.707000",
          "content": "<p>Hi, Denny, I use yolov3 (with 2 class filter) and have a LB score 2.3. I am trying with yolo5X now, but only got LB around 2.0. I guess yolo5x should achieve better performance than yolov3. Do you know any possible reasons? FYI, I already enable the \"multi-label\" for nms in yolov5. Thanks!    </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1233744,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-10T16:36:36.297000",
          "content": "<p>One more question, I'd like to try your strategy. How is \"ensemble models\" performed?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234302,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-03-11T05:48:14.987000",
          "content": "<p>From my testing, yolo5x is better at large image, did you resize your image to 1024? Plus I use yolo5x6, a variation of yolo5x. you can find weight file <a href=\"https://github.com/ultralytics/yolov5/releases\" target=\"_blank\">here</a> .</p>\n<p>Tried a few ensemble methods, the one works so far is removing the exact same bboxes(same class + IOU &gt; 0.95) then adding the rest bboxes together while keeps conf score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1234661,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-11T13:18:47.473000",
          "content": "<p>Cool, thanks! Will try your ensemble strategy. Yeah, I train yolo5x on 1024, and use the default augmentations. Are you using any EXTRA augmentations (to enlarge the dataset for yolov5) or label fusion (from 3 radiologists) for preprocessing? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1237315,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-03-14T03:39:24.453000",
          "content": "<p>Good luck. </p>\n<p>And about augmentation, I tried from very little to very heavy augmentations, heavy augmentations improved CV not LB, I have to turn down aug to improve LB score a little bit. I'm confused. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1238069,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-03-14T16:05:32.893000",
          "content": "<p>Thanks for all the feedback Denny, it has been helpful…. but for some reason our model performs quite terribly (LB Score 0.05x), which is hardly better than submitting a file of all no findings…. we check the boxes from the submission file back to the original images and they look reasonable… On the test data we looked at the predicted boxes literally fall right on top of most radiologists findings. We also got rid of overlapping boxes of the same class IOU &gt; 0.9 and somehow got the same score as before we removed them… We found a program to calculate mAP on the training set so we can at least see how it would score us there…. This is so frustrating! How can we be using the same model. I understand how to ensemble the models and do folds etc. but no reason to try any of that with the LB score we currently have. Will let everyone know what our mAP calculation on the training data comes out to… it should be high looking at the boxes, I really am not sure what is going on/to try next if anyone has suggestions…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1239624,
          "author_name": "Denny Wang",
          "author_url": "",
          "post_date": "2021-03-15T21:18:07.520000",
          "content": "<p>It's hard to tell what's missing in your case without seeing your notebook.</p>\n<p>For my experience, after 15 epochs, normally mAP:50 = ~0.4, mAP:0.5:0.95 = ~0.2, and LB = ~0.18 (without 2 classes filter)</p>\n<p>And if you specify val data in the data yaml,  Yolo script should be able to generate mAP:50 and mAP:0.5:0.95 and other metrics automatically. </p>\n<p>my yaml seen below</p>\n<p>==data.yaml==</p>\n<p>names:</p>\n<ul>\n<li>Aortic enlargement</li>\n<li>Atelectasis</li>\n<li>Calcification</li>\n<li>Cardiomegaly</li>\n<li>Consolidation</li>\n<li>ILD</li>\n<li>Infiltration</li>\n<li>Lung Opacity</li>\n<li>Nodule/Mass</li>\n<li>Other lesion</li>\n<li>Pleural effusion</li>\n<li>Pleural thickening</li>\n<li>Pneumothorax</li>\n<li>Pulmonary fibrosis<br>\nnc: 14<br>\ntrain: chest-abnormal/train.txt<br>\nval: chest-abnormal/val.txt</li>\n</ul>\n<p>==end==</p>\n<p>==val.txt==<br>\nchest-abnormal/yolo/images/val/78057413f8f7e8a3b1decc815e1f509e.png<br>\nchest-abnormal/yolo/images/val/d625684a437d0b9f622fec5329c6d7af.png<br>\nchest-abnormal/yolo/images/val/163898fbc57f00f58ad27e72031a541f.png<br>\nchest-abnormal/yolo/images/val/d9245845860a540560404f47bcda1716.png<br>\nchest-abnormal/yolo/images/val/7729fbb58b8006fe2b8305f9b2b33883.png<br>\nchest-abnormal/yolo/images/val/95b32ce2a10f57629eb63830376237ca.png<br>\nchest-abnormal/yolo/images/val/723f93a3f0e3905c498839d23b961631.png<br>\n…<br>\n==end==</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1215896,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-02-24T03:05:12.657000",
      "content": "<p>Hello,how much score can yolov5 get?</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1217226,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-02-24T21:18:45.853000",
          "content": "<p>I have not been able to make a submission file even though I have model trained, just getting the output in the right format so I can get back to you when we get a baseline score, but there are lots of people using this network, and it is one of the most successful objection detection models around these days. I would look at the leaderboard, if you noticed, Denny Wang who commented above is also using yolov5 and currently 16th on the LB… which is a score of 0.272</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1217313,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-24T23:59:10.137000",
          "content": "<p>Thank you, it is an efficient method.I do not have enough gpus.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1217376,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-02-25T01:59:03.283000",
          "content": "<p>Right, one can never have too many GPUs xD</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1217384,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-25T02:08:42.767000",
          "content": "<p>We can run it on Kaggle,single fold scored 0.15x on lb :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1218188,
          "author_name": "Daniel Hagan",
          "author_url": "",
          "post_date": "2021-02-25T16:24:31.423000",
          "content": "<p>Awesome, not too bad! We fixed our problem compiling the model so we should have a submission in of our own today too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1205750,
      "author_name": "Daniel Hagan",
      "author_url": "",
      "post_date": "2021-02-17T00:18:56.713000",
      "content": "<p><strong>ISSUE RESOLVED</strong></p>\n<p>Here is the output from PyCharm we print out the loss and the gradient (one from each GPU):<br>\n23/73 [========&gt;…………………] - ETA: 4:54 - loss: 24.2436loss =tf.Tensor(24.243551, shape=(), dtype=float32)<br>\ngrad: 44.53920408875909<br>\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)<br>\ngrad: 44.539202217558866<br>\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)<br>\n24/73 [========&gt;…………………] - ETA: 4:48 - loss: 23.7136loss =tf.Tensor(23.713629, shape=(), dtype=float32)<br>\ngrad: 61.30926377767098<br>\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)<br>\ngrad: 61.30925365406151<br>\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)<br>\n25/73 [=========&gt;………………..] - ETA: 4:40 - loss: 23.2499loss =tf.Tensor(23.249947, shape=(), dtype=float32)<br>\ngrad: 1.3752864851022184<br>\nloss: tf.Tensor(1.1149328, shape=(), dtype=float32)<br>\ngrad: 56.96218364695795<br>\nloss: tf.Tensor(0.97858346, shape=(), dtype=float32)<br>\n26/73 [=========&gt;………………..] - ETA: 4:28 - loss: 2.0935 loss =tf.Tensor(2.0935163, shape=(), dtype=float32)</p>\n<p>Process finished with exit code 0</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1205738": "Currently we are having issues training a YOLOv5 (using Adam optimizer) network on the data. Even with shuffling, it will only train up to step 26 (if you use batch size = 180 otherwise it will be step 51 with batch size = 90) where it unexpectedly stops, no error message is given as if the program ran successfully... Is anyone else having a similar issue? Here is a link to the code we started with[ https://github.com/jahongir7174/YOLOv5-tf](url)",
    "1210789": "Did you install and configure wandb properly? \n\nWe trained YOLOv5 on Colab without problem.\n==training==\n`!python train.py --img 512 --batch 16 --epochs 60 --data chest-abnormal/vinbigdata.yaml --weights yolov5x.pt --cache `\n==vinbigdata.yaml==,\n\nyaml:\nnames:\n- Aortic enlargement\n- Atelectasis\n- Calcification\n- Cardiomegaly\n- Consolidation\n- ILD\n- Infiltration\n- Lung Opacity\n- Nodule/Mass\n- Other lesion\n- Pleural effusion\n- Pleural thickening\n- Pneumothorax\n- Pulmonary fibrosis\nnc: 14\ntrain: chest-abnormal/train.txt\nval: chest-abnormal/val.txt\n\n\n==log==\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     57/59     9.82G   0.03037    0.0175   0.01142   0.05929        25       512: 100%|██████████| 220/220 [01:02<00:00,  3.54it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10<00:00,  5.11it/s]\n                 all         879    7.22e+03       0.403       0.273       0.232       0.101\n\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     58/59     9.82G   0.03001   0.01708   0.01122   0.05831        36       512: 100%|██████████| 220/220 [01:02<00:00,  3.53it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:10<00:00,  5.14it/s]\n                 all         879    7.22e+03       0.393       0.275        0.23       0.103\n\n     Epoch   gpu_mem       box       obj       cls     total   targets  img_size\n     59/59     9.82G   0.02999   0.01684   0.01104   0.05787        52       512: 100%|██████████| 220/220 [01:02<00:00,  3.53it/s]\n               Class      Images     Targets           P           R      mAP@.5  mAP@.5:.95: 100%|██████████| 55/55 [00:13<00:00,  3.97it/s]\n",
    "1215896": "Hello,how much score can yolov5 get?",
    "1205750": "**ISSUE RESOLVED**\n\nHere is the output from PyCharm we print out the loss and the gradient (one from each GPU):\n23/73 [========>.....................] - ETA: 4:54 - loss: 24.2436loss =tf.Tensor(24.243551, shape=(), dtype=float32)\ngrad: 44.53920408875909\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)\ngrad: 44.539202217558866\nloss: tf.Tensor(11.856814, shape=(), dtype=float32)\n24/73 [========>.....................] - ETA: 4:48 - loss: 23.7136loss =tf.Tensor(23.713629, shape=(), dtype=float32)\ngrad: 61.30926377767098\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)\ngrad: 61.30925365406151\nloss: tf.Tensor(11.624973, shape=(), dtype=float32)\n25/73 [=========>....................] - ETA: 4:40 - loss: 23.2499loss =tf.Tensor(23.249947, shape=(), dtype=float32)\ngrad: 1.3752864851022184\nloss: tf.Tensor(1.1149328, shape=(), dtype=float32)\ngrad: 56.96218364695795\nloss: tf.Tensor(0.97858346, shape=(), dtype=float32)\n26/73 [=========>....................] - ETA: 4:28 - loss: 2.0935 loss =tf.Tensor(2.0935163, shape=(), dtype=float32)\n\nProcess finished with exit code 0"
  }
}