{
  "id": 229636,
  "title": "Part of 3rd Place Sol: Multi-label Classifier-based PP",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/229636",
  "author_name": "ShinSiang",
  "post_date": "2021-03-31T03:51:21.413000",
  "votes": 39,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Codes: <a href=\"https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021\" target=\"_blank\">https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021</a></p>\n<p>We would like to thank the competition host(s) and Kaggle for organizing the competition, and congratulate all the winners, and anyone who benefited in some way from the competition. Special thanks to my teammates <a href=\"https://www.kaggle.com/erniechiew\" target=\"_blank\">@erniechiew</a>, <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a>, <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a>, and <a href=\"https://www.kaggle.com/stephkua\" target=\"_blank\">@stephkua</a> who carried me in the competition especially for the detector part.</p>\n<p>This is a summary of our <strong>Multi-label Classifier-based Post-Processing (PP)</strong>. Another part of the solution (for the detector etc.) can be found <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229769\" target=\"_blank\">here</a>.</p>\n<h2>PP Suggested in Public Notebooks</h2>\n<p>I believe most participants are using 2-class classifier(s) for PP. In most public notebooks (e.g. <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">this notebook</a>), when the 2-class classifier confidence is below a certain threshold, the predicted boxes are removed.</p>\n<h2>Multiplication of Classifier and Detector Confidences</h2>\n<p>We tried a different way instead of just removing the boxes, i.e. if the classifier is confident that the image is a normal image, we scale down the box confidences by multiplying them with the classifier confidence.</p>\n<h2>Our 2-Class Classifier</h2>\n<ul>\n<li>Model: EfficientNet-B6, 5 folds</li>\n<li>Image Size: 512 x 512</li>\n<li>Augmentation:</li>\n</ul>\n<ol>\n<li>HorizontalFlip(p = 0.5)</li>\n<li>ShiftScaleRotate(rotate_limit = 10),</li>\n<li>RandomBrightnessContrast(p = 0.2),</li>\n<li>CLAHE(clip_limit=4.0, tile_grid_size=(8, 8), p = 0.5),</li>\n<li>Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))</li>\n</ol>\n<ul>\n<li>Optimizer: Adam, BS = 4, BCE loss, LR=0.001, ReduceLROnPlateau + EarlyStopping </li>\n<li>Val AUC: 0.9868</li>\n</ul>\n<p>We found out that this classifier actually performs worse than <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">this notebook</a> on the public LB. We did not further try another 2-class classifier model. By applying the notebook’s classifier predictions on our final detector (also following its threshold), the scores are as follows:<br>\nRemove Predicted Boxes: Public=0.315, Private=0.292<br>\nMultiplication with 2-class-filter score: Public=0.319, Private=0.295</p>\n<h2>Our Multi-label Classifiers</h2>\n<p>We also tested a multi-label classifier. The Intuitions are as follows:</p>\n<ol>\n<li>All these classes interact with each other, a multi-label classifier may be able to capture their interactions and learn better. </li>\n<li>Is it better if we could do class-independent PP for each class?</li>\n</ol>\n<p>Model Details:</p>\n<ul>\n<li>Model: EfficientNet-B4, 5 folds</li>\n<li>Labels: 15 classes, A class of an image is labelled 1/0.67/0.33 if 3/2/1 radiologists annotated the class on the image.</li>\n<li>Image Size: 1024x1024, 1024x1024_with_fixed_aspect_ratio, 1280x1280_with_fixed_aspect_ratio (weighted average of 3 models)</li>\n<li>Augmentation:</li>\n</ul>\n<ol>\n<li>HorizontalFlip(p=0.5),</li>\n<li>ShiftScaleRotate(scale_limit = 0.15, rotate_limit = 10, p = 0.5),</li>\n<li>RandomBrightnessContrast(p=0.5),</li>\n<li>Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))</li>\n</ol>\n<ul>\n<li>Optimizer: SGD(momentum=0.9), BS = 2, BCE loss for each class, LR=0.001, ReduceLROnPlateau + EarlyStopping </li>\n<li>Class14 AUC: <br>\n1024x1024: 0.9923<br>\n1024x1024_with_fixed_aspect_ratio: 0.9937<br>\n1280x1280_with_fixed_aspect_ratio: 0.9940</li>\n</ul>\n<p>Worthy mentioning that training with multi labels could also improve class14 AUC. We tried 2 multiplication-based PP methods using this classifier:</p>\n<ol>\n<li>Multiplication with classifier score of respective class: If the classifier score of the class is less than a threshold, the box confidences of the class are scaled down by multiplying the classifier score of the class. We also implemented class-independent thresholds based on the validation AUCs of the classes, i.e. classifier scores of classes with higher validation AUC are allowed to multiply the box confidences of more images.</li>\n<li>Multiplication with class14 score only: This is similar to the 2-class classifier but note that the multi-label training has also improved the class14 performance.</li>\n</ol>\n<p>Results:<br>\nMultiplication with classifier score of respective class: Public=0.348, Private=0.305<br>\nMultiplication with class14 score only: Public=0.332, Private=0.305</p>\n<p>“Multiplication with classifier score of respective class” improved public LB by a lot, hence we used it in a final submission. Unfortunately, the improvement does not generalize to the private LB. We are lucky that our rank remains without that improvement :D</p>",
  "messages": [
    {
      "id": 1257712,
      "postDate": "2021-03-31T03:51:21.413Z",
      "content": "<p>Codes: <a href=\"https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021\" target=\"_blank\">https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021</a></p>\n<p>We would like to thank the competition host(s) and Kaggle for organizing the competition, and congratulate all the winners, and anyone who benefited in some way from the competition. Special thanks to my teammates <a href=\"https://www.kaggle.com/erniechiew\" target=\"_blank\">@erniechiew</a>, <a href=\"https://www.kaggle.com/scusywxy\" target=\"_blank\">@scusywxy</a>, <a href=\"https://www.kaggle.com/zehuigong\" target=\"_blank\">@zehuigong</a>, and <a href=\"https://www.kaggle.com/stephkua\" target=\"_blank\">@stephkua</a> who carried me in the competition especially for the detector part.</p>\n<p>This is a summary of our <strong>Multi-label Classifier-based Post-Processing (PP)</strong>. Another part of the solution (for the detector etc.) can be found <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229769\" target=\"_blank\">here</a>.</p>\n<h2>PP Suggested in Public Notebooks</h2>\n<p>I believe most participants are using 2-class classifier(s) for PP. In most public notebooks (e.g. <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">this notebook</a>), when the 2-class classifier confidence is below a certain threshold, the predicted boxes are removed.</p>\n<h2>Multiplication of Classifier and Detector Confidences</h2>\n<p>We tried a different way instead of just removing the boxes, i.e. if the classifier is confident that the image is a normal image, we scale down the box confidences by multiplying them with the classifier confidence.</p>\n<h2>Our 2-Class Classifier</h2>\n<ul>\n<li>Model: EfficientNet-B6, 5 folds</li>\n<li>Image Size: 512 x 512</li>\n<li>Augmentation:</li>\n</ul>\n<ol>\n<li>HorizontalFlip(p = 0.5)</li>\n<li>ShiftScaleRotate(rotate_limit = 10),</li>\n<li>RandomBrightnessContrast(p = 0.2),</li>\n<li>CLAHE(clip_limit=4.0, tile_grid_size=(8, 8), p = 0.5),</li>\n<li>Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))</li>\n</ol>\n<ul>\n<li>Optimizer: Adam, BS = 4, BCE loss, LR=0.001, ReduceLROnPlateau + EarlyStopping </li>\n<li>Val AUC: 0.9868</li>\n</ul>\n<p>We found out that this classifier actually performs worse than <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">this notebook</a> on the public LB. We did not further try another 2-class classifier model. By applying the notebook’s classifier predictions on our final detector (also following its threshold), the scores are as follows:<br>\nRemove Predicted Boxes: Public=0.315, Private=0.292<br>\nMultiplication with 2-class-filter score: Public=0.319, Private=0.295</p>\n<h2>Our Multi-label Classifiers</h2>\n<p>We also tested a multi-label classifier. The Intuitions are as follows:</p>\n<ol>\n<li>All these classes interact with each other, a multi-label classifier may be able to capture their interactions and learn better. </li>\n<li>Is it better if we could do class-independent PP for each class?</li>\n</ol>\n<p>Model Details:</p>\n<ul>\n<li>Model: EfficientNet-B4, 5 folds</li>\n<li>Labels: 15 classes, A class of an image is labelled 1/0.67/0.33 if 3/2/1 radiologists annotated the class on the image.</li>\n<li>Image Size: 1024x1024, 1024x1024_with_fixed_aspect_ratio, 1280x1280_with_fixed_aspect_ratio (weighted average of 3 models)</li>\n<li>Augmentation:</li>\n</ul>\n<ol>\n<li>HorizontalFlip(p=0.5),</li>\n<li>ShiftScaleRotate(scale_limit = 0.15, rotate_limit = 10, p = 0.5),</li>\n<li>RandomBrightnessContrast(p=0.5),</li>\n<li>Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))</li>\n</ol>\n<ul>\n<li>Optimizer: SGD(momentum=0.9), BS = 2, BCE loss for each class, LR=0.001, ReduceLROnPlateau + EarlyStopping </li>\n<li>Class14 AUC: <br>\n1024x1024: 0.9923<br>\n1024x1024_with_fixed_aspect_ratio: 0.9937<br>\n1280x1280_with_fixed_aspect_ratio: 0.9940</li>\n</ul>\n<p>Worthy mentioning that training with multi labels could also improve class14 AUC. We tried 2 multiplication-based PP methods using this classifier:</p>\n<ol>\n<li>Multiplication with classifier score of respective class: If the classifier score of the class is less than a threshold, the box confidences of the class are scaled down by multiplying the classifier score of the class. We also implemented class-independent thresholds based on the validation AUCs of the classes, i.e. classifier scores of classes with higher validation AUC are allowed to multiply the box confidences of more images.</li>\n<li>Multiplication with class14 score only: This is similar to the 2-class classifier but note that the multi-label training has also improved the class14 performance.</li>\n</ol>\n<p>Results:<br>\nMultiplication with classifier score of respective class: Public=0.348, Private=0.305<br>\nMultiplication with class14 score only: Public=0.332, Private=0.305</p>\n<p>“Multiplication with classifier score of respective class” improved public LB by a lot, hence we used it in a final submission. Unfortunately, the improvement does not generalize to the private LB. We are lucky that our rank remains without that improvement :D</p>",
      "rawMarkdown": "Codes: https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021\n\nWe would like to thank the competition host(s) and Kaggle for organizing the competition, and congratulate all the winners, and anyone who benefited in some way from the competition. Special thanks to my teammates @erniechiew, @scusywxy, @zehuigong, and @stephkua who carried me in the competition especially for the detector part.\n\nThis is a summary of our **Multi-label Classifier-based Post-Processing (PP)**. Another part of the solution (for the detector etc.) can be found [here](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229769).\n\n## PP Suggested in Public Notebooks\nI believe most participants are using 2-class classifier(s) for PP. In most public notebooks (e.g. [this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter)), when the 2-class classifier confidence is below a certain threshold, the predicted boxes are removed.\n\n## Multiplication of Classifier and Detector Confidences\nWe tried a different way instead of just removing the boxes, i.e. if the classifier is confident that the image is a normal image, we scale down the box confidences by multiplying them with the classifier confidence.\n\n## Our 2-Class Classifier\n- Model: EfficientNet-B6, 5 folds\n- Image Size: 512 x 512\n- Augmentation:\n1. HorizontalFlip(p = 0.5)\n2. ShiftScaleRotate(rotate_limit = 10),\n3. RandomBrightnessContrast(p = 0.2),\n4. CLAHE(clip_limit=4.0, tile_grid_size=(8, 8), p = 0.5),\n5. Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))\n- Optimizer: Adam, BS = 4, BCE loss, LR=0.001, ReduceLROnPlateau + EarlyStopping \n- Val AUC: 0.9868\n\nWe found out that this classifier actually performs worse than [this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter) on the public LB. We did not further try another 2-class classifier model. By applying the notebook’s classifier predictions on our final detector (also following its threshold), the scores are as follows:\nRemove Predicted Boxes: Public=0.315, Private=0.292\nMultiplication with 2-class-filter score: Public=0.319, Private=0.295\n\n## Our Multi-label Classifiers\nWe also tested a multi-label classifier. The Intuitions are as follows:\n1. All these classes interact with each other, a multi-label classifier may be able to capture their interactions and learn better. \n2. Is it better if we could do class-independent PP for each class?\n\nModel Details:\n- Model: EfficientNet-B4, 5 folds\n- Labels: 15 classes, A class of an image is labelled 1/0.67/0.33 if 3/2/1 radiologists annotated the class on the image.\n- Image Size: 1024x1024, 1024x1024_with_fixed_aspect_ratio, 1280x1280_with_fixed_aspect_ratio (weighted average of 3 models)\n- Augmentation:\n1. HorizontalFlip(p=0.5),\n2. ShiftScaleRotate(scale_limit = 0.15, rotate_limit = 10, p = 0.5),\n3. RandomBrightnessContrast(p=0.5),\n4. Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))\n- Optimizer: SGD(momentum=0.9), BS = 2, BCE loss for each class, LR=0.001, ReduceLROnPlateau + EarlyStopping \n- Class14 AUC: \n1024x1024: 0.9923\n1024x1024_with_fixed_aspect_ratio: 0.9937\n1280x1280_with_fixed_aspect_ratio: 0.9940\n\nWorthy mentioning that training with multi labels could also improve class14 AUC. We tried 2 multiplication-based PP methods using this classifier:\n1. Multiplication with classifier score of respective class: If the classifier score of the class is less than a threshold, the box confidences of the class are scaled down by multiplying the classifier score of the class. We also implemented class-independent thresholds based on the validation AUCs of the classes, i.e. classifier scores of classes with higher validation AUC are allowed to multiply the box confidences of more images.\n2. Multiplication with class14 score only: This is similar to the 2-class classifier but note that the multi-label training has also improved the class14 performance.\n\nResults:\nMultiplication with classifier score of respective class: Public=0.348, Private=0.305\nMultiplication with class14 score only: Public=0.332, Private=0.305\n\n“Multiplication with classifier score of respective class” improved public LB by a lot, hence we used it in a final submission. Unfortunately, the improvement does not generalize to the private LB. We are lucky that our rank remains without that improvement :D\n\n\n",
      "votes": 39
    },
    {
      "id": 1257727,
      "postDate": "2021-03-31T04:05:32.370Z",
      "content": "<p>Nice, we did try the multilabel classification post-processing and the way we applied was pretty much similar to yours. But we got confused when it came to choosing the threshold for each class to decide whether to multiply the classifier score, we used hyperopt for tuning on validsets but it did not seem to have the unified results between my CV OOF and also the PB. Yours was at least having an improvement on PB. It was very nice hearing the similar idea</p>",
      "rawMarkdown": "Nice, we did try the multilabel classification post-processing and the way we applied was pretty much similar to yours. But we got confused when it came to choosing the threshold for each class to decide whether to multiply the classifier score, we used hyperopt for tuning on validsets but it did not seem to have the unified results between my CV OOF and also the PB. Yours was at least having an improvement on PB. It was very nice hearing the similar idea",
      "votes": 3,
      "replies": [
        {
          "id": 1257729,
          "postDate": "2021-03-31T04:13:51.133Z",
          "content": "<p>The optimal thresholds for CV and Public LB are pretty far off in our case too. In the end we decided to follow the public notebook I mentioned above. Its threshold covers 1912/3000 images, we used this number as a base. </p>",
          "rawMarkdown": "The optimal thresholds for CV and Public LB are pretty far off in our case too. In the end we decided to follow the public notebook I mentioned above. Its threshold covers 1912/3000 images, we used this number as a base. ",
          "votes": 2
        },
        {
          "id": 1257743,
          "postDate": "2021-03-31T04:32:10.157Z",
          "content": "<p>For the multi-label classifier, you relied on your valid AUC of each class to determine the thresholds. Can I know more detail about that ?</p>",
          "rawMarkdown": "For the multi-label classifier, you relied on your valid AUC of each class to determine the thresholds. Can I know more detail about that ?",
          "votes": 1
        },
        {
          "id": 1257776,
          "postDate": "2021-03-31T05:00:19.660Z",
          "content": "<p>Each class (excluding class14) have a rank_threshold based on this formula: np.round(aucs[i]/aucs.sum() x 1912 x 14)</p>\n<p>As I mentioned earlier, 1912 is the base number and I used it in this way. The public leaderboard improvement is still quite significant even if all classes are PPed with a same threshold.</p>",
          "rawMarkdown": "Each class (excluding class14) have a rank_threshold based on this formula: np.round(aucs[i]/aucs.sum() x 1912 x 14)\n\nAs I mentioned earlier, 1912 is the base number and I used it in this way. The public leaderboard improvement is still quite significant even if all classes are PPed with a same threshold.",
          "votes": 2
        },
        {
          "id": 1257818,
          "postDate": "2021-03-31T05:56:16.810Z",
          "content": "<p>If you don't mind, can I have a look at valid AUC score of each of your class? I want to compare your AUCs with mine</p>",
          "rawMarkdown": "If you don't mind, can I have a look at valid AUC score of each of your class? I want to compare your AUCs with mine",
          "votes": 1
        },
        {
          "id": 1257821,
          "postDate": "2021-03-31T05:59:16.507Z",
          "content": "<p>class 0 - 13 weighted average AUC of 3 models:<br>\n[0.97586354, 0.97834093, 0.96728132, 0.98092026, 0.98604685,<br>\n0.98757058, 0.9808205 , 0.96089901, 0.96671836, 0.91450653,<br>\n0.99094985, 0.95365022, 0.99780411, 0.97474175]</p>",
          "rawMarkdown": "class 0 - 13 weighted average AUC of 3 models:\n[0.97586354, 0.97834093, 0.96728132, 0.98092026, 0.98604685,\n0.98757058, 0.9808205 , 0.96089901, 0.96671836, 0.91450653,\n0.99094985, 0.95365022, 0.99780411, 0.97474175]",
          "votes": 2
        },
        {
          "id": 1257930,
          "postDate": "2021-03-31T08:12:40.587Z",
          "content": "<p>I got it. An image is No finding when ALL the values of the 14 positive classes is 0. <br>\nYou want to the multilabel classifier to agree on the 1912 No finding results with the 2 class-filter. </p>\n<p>So for each class id i from 0 to 13, you rank all the predictions by the class_id probability ascending, and pick up the class i probability at 1912th as the threshold -&gt; use this to get the 1912 predictions with lowest score to be the No-findings of class i. We do the same thing to the other 13 classes in the expectation of an \"ideal\" case where all the No-findings of class 0 -&gt; 13 agree with each other and we will have 1912 No findings as using the 2 class filter (But not in practice). You further let the classes with higher AUC to choose more images to be No findings by the above equation, with the total sum = 1912x14.</p>\n<p>That is a lot of thought. Thank you for the explanation.</p>",
          "rawMarkdown": "I got it. An image is No finding when ALL the values of the 14 positive classes is 0. \nYou want to the multilabel classifier to agree on the 1912 No finding results with the 2 class-filter. \n\nSo for each class id i from 0 to 13, you rank all the predictions by the class_id probability ascending, and pick up the class i probability at 1912th as the threshold -> use this to get the 1912 predictions with lowest score to be the No-findings of class i. We do the same thing to the other 13 classes in the expectation of an \"ideal\" case where all the No-findings of class 0 -> 13 agree with each other and we will have 1912 No findings as using the 2 class filter (But not in practice). You further let the classes with higher AUC to choose more images to be No findings by the above equation, with the total sum = 1912x14.\n\nThat is a lot of thought. Thank you for the explanation.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1257716,
      "postDate": "2021-03-31T03:58:04.753Z",
      "content": "<p>Congrats on 3rd place <a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> and team and thanks for sharing details solution</p>",
      "rawMarkdown": "Congrats on 3rd place @css919 and team and thanks for sharing details solution",
      "votes": 4
    },
    {
      "id": 1258134,
      "postDate": "2021-03-31T11:29:47.697Z",
      "content": "<p>Congratulations for 3rd place who were our Worth Competitor! <br>\nWe trained a Multi Label Classification Network . But i didn't improved our score , I decreased in our case. 😄</p>",
      "rawMarkdown": "Congratulations for 3rd place who were our Worth Competitor! \nWe trained a Multi Label Classification Network . But i didn't improved our score , I decreased in our case. 😄",
      "votes": 1
    },
    {
      "id": 1257920,
      "postDate": "2021-03-31T08:05:40.440Z",
      "content": "<p>Congratulations for 3rd place! I am expecting detection part of you and thanks for sharing your solution &lt;3 </p>",
      "rawMarkdown": "Congratulations for 3rd place! I am expecting detection part of you and thanks for sharing your solution <3 ",
      "votes": 2
    },
    {
      "id": 1257816,
      "postDate": "2021-03-31T05:51:39.253Z",
      "content": "<p>Thanks for sharing. I like your pp idea.</p>\n<p>I have a question here.</p>\n<p>About 15 classes, if 3 radiologists annotates that the image has a \"class a\", then the image's \"class a\" become active as 1? For example, in an image, when 3 doctors annotate \"class 0\" and 2 doctors annotate \"class 1\", is the class of the image as follows \"1, 0.67, 0, 0, 0, 0, …, 0\"?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Thanks for sharing. I like your pp idea.\n\n I have a question here.\n\nAbout 15 classes, if 3 radiologists annotates that the image has a \"class a\", then the image's \"class a\" become active as 1? For example, in an image, when 3 doctors annotate \"class 0\" and 2 doctors annotate \"class 1\", is the class of the image as follows \"1, 0.67, 0, 0, 0, 0, ..., 0\"?\n\nThanks.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1257819,
          "postDate": "2021-03-31T05:56:23.443Z",
          "content": "<p>Yes, this is exactly what we did.</p>",
          "rawMarkdown": "Yes, this is exactly what we did.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1257809,
      "postDate": "2021-03-31T05:44:32.927Z",
      "content": "<p>Congrats for your 3rd finish. It's was so fun to know your approaches</p>",
      "rawMarkdown": "Congrats for your 3rd finish. It's was so fun to know your approaches",
      "votes": 2
    },
    {
      "id": 1257720,
      "postDate": "2021-03-31T04:02:42.477Z",
      "content": "<p><a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> Congratulations on 3 rd Place and Thanks for sharing the approach</p>",
      "rawMarkdown": "@css919 Congratulations on 3 rd Place and Thanks for sharing the approach",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1257727,
      "author_name": "Liam Nguyen",
      "author_url": "",
      "post_date": "2021-03-31T04:05:32.370000",
      "content": "<p>Nice, we did try the multilabel classification post-processing and the way we applied was pretty much similar to yours. But we got confused when it came to choosing the threshold for each class to decide whether to multiply the classifier score, we used hyperopt for tuning on validsets but it did not seem to have the unified results between my CV OOF and also the PB. Yours was at least having an improvement on PB. It was very nice hearing the similar idea</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1257729,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-03-31T04:13:51.133000",
          "content": "<p>The optimal thresholds for CV and Public LB are pretty far off in our case too. In the end we decided to follow the public notebook I mentioned above. Its threshold covers 1912/3000 images, we used this number as a base. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1257743,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-03-31T04:32:10.157000",
          "content": "<p>For the multi-label classifier, you relied on your valid AUC of each class to determine the thresholds. Can I know more detail about that ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1257776,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-03-31T05:00:19.660000",
          "content": "<p>Each class (excluding class14) have a rank_threshold based on this formula: np.round(aucs[i]/aucs.sum() x 1912 x 14)</p>\n<p>As I mentioned earlier, 1912 is the base number and I used it in this way. The public leaderboard improvement is still quite significant even if all classes are PPed with a same threshold.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1257818,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-03-31T05:56:16.810000",
          "content": "<p>If you don't mind, can I have a look at valid AUC score of each of your class? I want to compare your AUCs with mine</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1257821,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-03-31T05:59:16.507000",
          "content": "<p>class 0 - 13 weighted average AUC of 3 models:<br>\n[0.97586354, 0.97834093, 0.96728132, 0.98092026, 0.98604685,<br>\n0.98757058, 0.9808205 , 0.96089901, 0.96671836, 0.91450653,<br>\n0.99094985, 0.95365022, 0.99780411, 0.97474175]</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1257930,
          "author_name": "Liam Nguyen",
          "author_url": "",
          "post_date": "2021-03-31T08:12:40.587000",
          "content": "<p>I got it. An image is No finding when ALL the values of the 14 positive classes is 0. <br>\nYou want to the multilabel classifier to agree on the 1912 No finding results with the 2 class-filter. </p>\n<p>So for each class id i from 0 to 13, you rank all the predictions by the class_id probability ascending, and pick up the class i probability at 1912th as the threshold -&gt; use this to get the 1912 predictions with lowest score to be the No-findings of class i. We do the same thing to the other 13 classes in the expectation of an \"ideal\" case where all the No-findings of class 0 -&gt; 13 agree with each other and we will have 1912 No findings as using the 2 class filter (But not in practice). You further let the classes with higher AUC to choose more images to be No findings by the above equation, with the total sum = 1912x14.</p>\n<p>That is a lot of thought. Thank you for the explanation.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1257716,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-31T03:58:04.753000",
      "content": "<p>Congrats on 3rd place <a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> and team and thanks for sharing details solution</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1258134,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-03-31T11:29:47.697000",
      "content": "<p>Congratulations for 3rd place who were our Worth Competitor! <br>\nWe trained a Multi Label Classification Network . But i didn't improved our score , I decreased in our case. 😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1257920,
      "author_name": "HungNT",
      "author_url": "",
      "post_date": "2021-03-31T08:05:40.440000",
      "content": "<p>Congratulations for 3rd place! I am expecting detection part of you and thanks for sharing your solution &lt;3 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1257816,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2021-03-31T05:51:39.253000",
      "content": "<p>Thanks for sharing. I like your pp idea.</p>\n<p>I have a question here.</p>\n<p>About 15 classes, if 3 radiologists annotates that the image has a \"class a\", then the image's \"class a\" become active as 1? For example, in an image, when 3 doctors annotate \"class 0\" and 2 doctors annotate \"class 1\", is the class of the image as follows \"1, 0.67, 0, 0, 0, 0, …, 0\"?</p>\n<p>Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1257819,
          "author_name": "ShinSiang",
          "author_url": "",
          "post_date": "2021-03-31T05:56:23.443000",
          "content": "<p>Yes, this is exactly what we did.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1257809,
      "author_name": "Binh Nguyen",
      "author_url": "",
      "post_date": "2021-03-31T05:44:32.927000",
      "content": "<p>Congrats for your 3rd finish. It's was so fun to know your approaches</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1257720,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-31T04:02:42.477000",
      "content": "<p><a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> Congratulations on 3 rd Place and Thanks for sharing the approach</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1257712": "Codes: https://github.com/shinsiangchoong/kaggle_vinbigdata_classifierPP_2021\n\nWe would like to thank the competition host(s) and Kaggle for organizing the competition, and congratulate all the winners, and anyone who benefited in some way from the competition. Special thanks to my teammates @erniechiew, @scusywxy, @zehuigong, and @stephkua who carried me in the competition especially for the detector part.\n\nThis is a summary of our **Multi-label Classifier-based Post-Processing (PP)**. Another part of the solution (for the detector etc.) can be found [here](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229769).\n\n## PP Suggested in Public Notebooks\nI believe most participants are using 2-class classifier(s) for PP. In most public notebooks (e.g. [this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter)), when the 2-class classifier confidence is below a certain threshold, the predicted boxes are removed.\n\n## Multiplication of Classifier and Detector Confidences\nWe tried a different way instead of just removing the boxes, i.e. if the classifier is confident that the image is a normal image, we scale down the box confidences by multiplying them with the classifier confidence.\n\n## Our 2-Class Classifier\n- Model: EfficientNet-B6, 5 folds\n- Image Size: 512 x 512\n- Augmentation:\n1. HorizontalFlip(p = 0.5)\n2. ShiftScaleRotate(rotate_limit = 10),\n3. RandomBrightnessContrast(p = 0.2),\n4. CLAHE(clip_limit=4.0, tile_grid_size=(8, 8), p = 0.5),\n5. Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))\n- Optimizer: Adam, BS = 4, BCE loss, LR=0.001, ReduceLROnPlateau + EarlyStopping \n- Val AUC: 0.9868\n\nWe found out that this classifier actually performs worse than [this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter) on the public LB. We did not further try another 2-class classifier model. By applying the notebook’s classifier predictions on our final detector (also following its threshold), the scores are as follows:\nRemove Predicted Boxes: Public=0.315, Private=0.292\nMultiplication with 2-class-filter score: Public=0.319, Private=0.295\n\n## Our Multi-label Classifiers\nWe also tested a multi-label classifier. The Intuitions are as follows:\n1. All these classes interact with each other, a multi-label classifier may be able to capture their interactions and learn better. \n2. Is it better if we could do class-independent PP for each class?\n\nModel Details:\n- Model: EfficientNet-B4, 5 folds\n- Labels: 15 classes, A class of an image is labelled 1/0.67/0.33 if 3/2/1 radiologists annotated the class on the image.\n- Image Size: 1024x1024, 1024x1024_with_fixed_aspect_ratio, 1280x1280_with_fixed_aspect_ratio (weighted average of 3 models)\n- Augmentation:\n1. HorizontalFlip(p=0.5),\n2. ShiftScaleRotate(scale_limit = 0.15, rotate_limit = 10, p = 0.5),\n3. RandomBrightnessContrast(p=0.5),\n4. Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225))\n- Optimizer: SGD(momentum=0.9), BS = 2, BCE loss for each class, LR=0.001, ReduceLROnPlateau + EarlyStopping \n- Class14 AUC: \n1024x1024: 0.9923\n1024x1024_with_fixed_aspect_ratio: 0.9937\n1280x1280_with_fixed_aspect_ratio: 0.9940\n\nWorthy mentioning that training with multi labels could also improve class14 AUC. We tried 2 multiplication-based PP methods using this classifier:\n1. Multiplication with classifier score of respective class: If the classifier score of the class is less than a threshold, the box confidences of the class are scaled down by multiplying the classifier score of the class. We also implemented class-independent thresholds based on the validation AUCs of the classes, i.e. classifier scores of classes with higher validation AUC are allowed to multiply the box confidences of more images.\n2. Multiplication with class14 score only: This is similar to the 2-class classifier but note that the multi-label training has also improved the class14 performance.\n\nResults:\nMultiplication with classifier score of respective class: Public=0.348, Private=0.305\nMultiplication with class14 score only: Public=0.332, Private=0.305\n\n“Multiplication with classifier score of respective class” improved public LB by a lot, hence we used it in a final submission. Unfortunately, the improvement does not generalize to the private LB. We are lucky that our rank remains without that improvement :D\n\n\n",
    "1257727": "Nice, we did try the multilabel classification post-processing and the way we applied was pretty much similar to yours. But we got confused when it came to choosing the threshold for each class to decide whether to multiply the classifier score, we used hyperopt for tuning on validsets but it did not seem to have the unified results between my CV OOF and also the PB. Yours was at least having an improvement on PB. It was very nice hearing the similar idea",
    "1257716": "Congrats on 3rd place @css919 and team and thanks for sharing details solution",
    "1258134": "Congratulations for 3rd place who were our Worth Competitor! \nWe trained a Multi Label Classification Network . But i didn't improved our score , I decreased in our case. 😄",
    "1257920": "Congratulations for 3rd place! I am expecting detection part of you and thanks for sharing your solution <3 ",
    "1257816": "Thanks for sharing. I like your pp idea.\n\n I have a question here.\n\nAbout 15 classes, if 3 radiologists annotates that the image has a \"class a\", then the image's \"class a\" become active as 1? For example, in an image, when 3 doctors annotate \"class 0\" and 2 doctors annotate \"class 1\", is the class of the image as follows \"1, 0.67, 0, 0, 0, 0, ..., 0\"?\n\nThanks.\n\n",
    "1257809": "Congrats for your 3rd finish. It's was so fun to know your approaches",
    "1257720": "@css919 Congratulations on 3 rd Place and Thanks for sharing the approach"
  }
}