{
  "id": 227425,
  "title": "Classification",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/227425",
  "author_name": "aqx",
  "post_date": "2021-03-20T12:43:10.625000",
  "votes": 8,
  "comment_count": 16,
  "views": 0,
  "content": "<p>My valid AUC for classification on single fold can't seem to break the 0.992. Was wondering what are some of the architecture people are using. I felt that a good 2-class classifier is pretty important (maybe im wrong) but there isn't much discussion on the classification stage</p>",
  "messages": [
    {
      "id": 1248009,
      "postDate": "2021-03-22T09:13:46.130Z",
      "content": "<p>Kindly remind you that the public test set account for ONLY 10% of the whole test dataset and that amount of data will 1) introduce a big shakeup with blind trust and no robust local cross-validation; 2) doesn't have enough samples to show the differences between 2 models that have 0.1% or 0.2% different in AUC, which is already a very skewed metric in opinion.</p>",
      "rawMarkdown": "Kindly remind you that the public test set account for ONLY 10% of the whole test dataset and that amount of data will 1) introduce a big shakeup with blind trust and no robust local cross-validation; 2) doesn't have enough samples to show the differences between 2 models that have 0.1% or 0.2% different in AUC, which is already a very skewed metric in opinion.",
      "votes": 7,
      "replies": [
        {
          "id": 1248036,
          "postDate": "2021-03-22T09:31:51.653Z",
          "content": "<p>Yes i agree totally with the need for a robust CV. </p>",
          "rawMarkdown": "Yes i agree totally with the need for a robust CV. ",
          "votes": 1
        },
        {
          "id": 1248053,
          "postDate": "2021-03-22T09:46:59.300Z",
          "content": "<p>I don't think robust CV is possible in this competition. </p>\n<ul>\n<li>Too few observations</li>\n<li>Too much label noise</li>\n<li>Bias in data collection</li>\n<li>Different and unknown train/test labeling process</li>\n</ul>\n<p>In my last five kaggle competitions (VinBigData, Rainforest, Lyft, Cornell, Covid) only Lyft was a competition where proper local CV was possible and even there you had to figure out a few sampling parameters to match the test set…</p>",
          "rawMarkdown": "I don't think robust CV is possible in this competition. \n- Too few observations\n- Too much label noise\n- Bias in data collection\n- Different and unknown train/test labeling process\n\n\nIn my last five kaggle competitions (VinBigData, Rainforest, Lyft, Cornell, Covid) only Lyft was a competition where proper local CV was possible and even there you had to figure out a few sampling parameters to match the test set...",
          "votes": 9
        },
        {
          "id": 1248134,
          "postDate": "2021-03-22T11:28:32.460Z",
          "content": "<p>I ran another little experiment on whether we can trust the public LB on the 2-class classifier: I used different shares of unhealthy images for the whole test set (3000 images) and simulated 100,000 times a random sampling of 300 images without replacement for each of these shares. Below you can see that the standard deviation of the share of unhealthy images in the public LB is higher than 2%. This is quite significant as I realized that the threshold used for my 2-class classifier ist very sensitive to small changes.. Worth noticing is also that the standard deviation increases with a higher share of unhealthy images in the whole test set.</p>\n<table>\n<thead>\n<tr>\n<th>share of unhealthy images (out of 3000)</th>\n<th>standard deviation in the public test set</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.2</td>\n<td>0.0219</td>\n</tr>\n<tr>\n<td>0.25</td>\n<td>0.0238</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.0251</td>\n</tr>\n<tr>\n<td>0.35</td>\n<td>0.0262</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "I ran another little experiment on whether we can trust the public LB on the 2-class classifier: I used different shares of unhealthy images for the whole test set (3000 images) and simulated 100,000 times a random sampling of 300 images without replacement for each of these shares. Below you can see that the standard deviation of the share of unhealthy images in the public LB is higher than 2%. This is quite significant as I realized that the threshold used for my 2-class classifier ist very sensitive to small changes.. Worth noticing is also that the standard deviation increases with a higher share of unhealthy images in the whole test set.\n| share of unhealthy images (out of 3000) | standard deviation in the public test set |\n| --- | --- |\n| 0.2 | 0.0219 |\n| 0.25 | 0.0238 |\n| 0.3 | 0.0251 |\n| 0.35 | 0.0262 |\n",
          "votes": 5
        },
        {
          "id": 1252583,
          "postDate": "2021-03-25T19:46:03.650Z",
          "content": "<p>Since we find pretty unstable in CV sometimes ,<br>\nWe can conclude the same problem as in Rainforest , where the first place solution ( <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> and team) determined their submission selection based on LB score.</p>",
          "rawMarkdown": "Since we find pretty unstable in CV sometimes ,\nWe can conclude the same problem as in Rainforest , where the first place solution ( @ilu000 and team) determined their submission selection based on LB score."
        }
      ]
    },
    {
      "id": 1246031,
      "postDate": "2021-03-20T12:43:10.627Z",
      "content": "<p>My valid AUC for classification on single fold can't seem to break the 0.992. Was wondering what are some of the architecture people are using. I felt that a good 2-class classifier is pretty important (maybe im wrong) but there isn't much discussion on the classification stage</p>",
      "rawMarkdown": "My valid AUC for classification on single fold can't seem to break the 0.992. Was wondering what are some of the architecture people are using. I felt that a good 2-class classifier is pretty important (maybe im wrong) but there isn't much discussion on the classification stage",
      "votes": 8
    },
    {
      "id": 1246297,
      "postDate": "2021-03-20T16:17:26.337Z",
      "content": "<p>Thanks for all the sharing! Seems like a big architecture is common in this classification stage. I achieved 0.992 with EffNetB4 trained on 1024-&gt;512. Using B5/B6 performs worst in my case, albeit the heavy augmentations used. Somehow the 2-class classifier provided <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">in this notebook</a> (AUC 0.98) gives me the best score in public LB.</p>",
      "rawMarkdown": "Thanks for all the sharing! Seems like a big architecture is common in this classification stage. I achieved 0.992 with EffNetB4 trained on 1024->512. Using B5/B6 performs worst in my case, albeit the heavy augmentations used. Somehow the 2-class classifier provided [in this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter) (AUC 0.98) gives me the best score in public LB.",
      "votes": 3,
      "replies": [
        {
          "id": 1246317,
          "postDate": "2021-03-20T16:36:04Z",
          "content": "<p>yeah, for me too, i used this <a href=\"https://www.kaggle.com/mrinath/2-class-classifier-pipeline-using-effnet\" target=\"_blank\">notebook</a> for train my own 2 class filter and get auc &gt; 0.99 in my 5 folds, but it is not better than the notebook you mentioned. ☹️</p>",
          "rawMarkdown": "yeah, for me too, i used this [notebook](https://www.kaggle.com/mrinath/2-class-classifier-pipeline-using-effnet) for train my own 2 class filter and get auc > 0.99 in my 5 folds, but it is not better than the notebook you mentioned. ☹️"
        },
        {
          "id": 1246328,
          "postDate": "2021-03-20T16:49:08.290Z",
          "content": "<p>You can try ensemble learning. I mixed the 2-class classifier provided in this <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">notebook</a> with my model to get better score.</p>",
          "rawMarkdown": "You can try ensemble learning. I mixed the 2-class classifier provided in this [notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter) with my model to get better score."
        },
        {
          "id": 1246329,
          "postDate": "2021-03-20T16:49:45.277Z",
          "content": "<p><a href=\"https://www.kaggle.com/angqx95\" target=\"_blank\">@angqx95</a> <a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> The same for me, that notebook gives me the best score. Have you guys been able find better parameters or are you using <br>\nlow_thr  = 0.08<br>\nhigh_thr = 0.95<br>\n?</p>",
          "rawMarkdown": "@angqx95 @adrielcabral The same for me, that notebook gives me the best score. Have you guys been able find better parameters or are you using \nlow_thr  = 0.08\nhigh_thr = 0.95\n?"
        },
        {
          "id": 1247623,
          "postDate": "2021-03-21T22:05:28.847Z",
          "content": "<p><a href=\"https://www.kaggle.com/linusj79\" target=\"_blank\">@linusj79</a> sorry, i just see now your comment, i use low_thr = 0.08 high_thr = 0.95</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> u can say how much your mixed 2 class improve lb ?</p>",
          "rawMarkdown": "@linusj79 sorry, i just see now your comment, i use low_thr = 0.08 high_thr = 0.95\n\n@h053473666 u can say how much your mixed 2 class improve lb ?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1246214,
      "postDate": "2021-03-20T14:44:58.017Z",
      "content": "<p>I also got 0.992 (efficientnet b8 with adversarial training).</p>",
      "rawMarkdown": "I also got 0.992 (efficientnet b8 with adversarial training).",
      "votes": 3
    },
    {
      "id": 1246163,
      "postDate": "2021-03-20T13:34:47.330Z",
      "content": "<p>efficientnet b0 224px - auc 0.987<br>\nefficientnet b5 456px - auc 0.992<br>\nefficientnet b5 600px - auc 0.993<br>\nefficientnet b6 528px - auc 0.992<br>\nresnet200d 600px - auc 0.992 (short training, I'm trying to train longer)</p>\n<p>These are results of my experiments.<br>\nI didn't use hard augmentations like cutmix.<br>\nAlso I tried train with concat dataset of VBD and NIH.</p>\n<p>And I'm going to</p>\n<ul>\n<li>train completely concat dataset of VBD and NIH</li>\n<li>increase image size</li>\n<li>try other architectures (nfnet, seresnet, etc ?)</li>\n</ul>",
      "rawMarkdown": "efficientnet b0 224px - auc 0.987\nefficientnet b5 456px - auc 0.992\nefficientnet b5 600px - auc 0.993\nefficientnet b6 528px - auc 0.992\nresnet200d 600px - auc 0.992 (short training, I'm trying to train longer)\n\nThese are results of my experiments.\nI didn't use hard augmentations like cutmix.\nAlso I tried train with concat dataset of VBD and NIH.\n\nAnd I'm going to\n- train completely concat dataset of VBD and NIH\n- increase image size\n- try other architectures (nfnet, seresnet, etc ?)",
      "votes": 3,
      "replies": [
        {
          "id": 1246305,
          "postDate": "2021-03-20T16:19:58.163Z",
          "content": "<p>I didn't attempt cutmix as i felt that it could potentially retrieve a crop of an abnormal image that doesn't include any abnormality, thereby creating unnecessary noise.</p>",
          "rawMarkdown": "I didn't attempt cutmix as i felt that it could potentially retrieve a crop of an abnormal image that doesn't include any abnormality, thereby creating unnecessary noise.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1246547,
      "postDate": "2021-03-20T20:43:36.073Z",
      "content": "<p>Does the classification really help? For us, it doesn't help at all.</p>",
      "rawMarkdown": "Does the classification really help? For us, it doesn't help at all.",
      "votes": 1,
      "replies": [
        {
          "id": 1246689,
          "postDate": "2021-03-21T02:24:38.217Z",
          "content": "<p>You mean the score doesn't vary with different classifiers? Or you didnt have a complete separate 2class pipeline?</p>",
          "rawMarkdown": "You mean the score doesn't vary with different classifiers? Or you didnt have a complete separate 2class pipeline?"
        }
      ]
    },
    {
      "id": 1246234,
      "postDate": "2021-03-20T15:06:57.317Z",
      "content": "<p>I use Efficientnet b7(noisy-student) to get the best auc.</p>",
      "rawMarkdown": "I use Efficientnet b7(noisy-student) to get the best auc.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1248009,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2021-03-22T09:13:46.130000",
      "content": "<p>Kindly remind you that the public test set account for ONLY 10% of the whole test dataset and that amount of data will 1) introduce a big shakeup with blind trust and no robust local cross-validation; 2) doesn't have enough samples to show the differences between 2 models that have 0.1% or 0.2% different in AUC, which is already a very skewed metric in opinion.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1248036,
          "author_name": "aqx",
          "author_url": "",
          "post_date": "2021-03-22T09:31:51.653000",
          "content": "<p>Yes i agree totally with the need for a robust CV. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1248053,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2021-03-22T09:46:59.300000",
          "content": "<p>I don't think robust CV is possible in this competition. </p>\n<ul>\n<li>Too few observations</li>\n<li>Too much label noise</li>\n<li>Bias in data collection</li>\n<li>Different and unknown train/test labeling process</li>\n</ul>\n<p>In my last five kaggle competitions (VinBigData, Rainforest, Lyft, Cornell, Covid) only Lyft was a competition where proper local CV was possible and even there you had to figure out a few sampling parameters to match the test set…</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1248134,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-22T11:28:32.460000",
          "content": "<p>I ran another little experiment on whether we can trust the public LB on the 2-class classifier: I used different shares of unhealthy images for the whole test set (3000 images) and simulated 100,000 times a random sampling of 300 images without replacement for each of these shares. Below you can see that the standard deviation of the share of unhealthy images in the public LB is higher than 2%. This is quite significant as I realized that the threshold used for my 2-class classifier ist very sensitive to small changes.. Worth noticing is also that the standard deviation increases with a higher share of unhealthy images in the whole test set.</p>\n<table>\n<thead>\n<tr>\n<th>share of unhealthy images (out of 3000)</th>\n<th>standard deviation in the public test set</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.2</td>\n<td>0.0219</td>\n</tr>\n<tr>\n<td>0.25</td>\n<td>0.0238</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.0251</td>\n</tr>\n<tr>\n<td>0.35</td>\n<td>0.0262</td>\n</tr>\n</tbody>\n</table>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1252583,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-25T19:46:03.650000",
          "content": "<p>Since we find pretty unstable in CV sometimes ,<br>\nWe can conclude the same problem as in Rainforest , where the first place solution ( <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> and team) determined their submission selection based on LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1246297,
      "author_name": "aqx",
      "author_url": "",
      "post_date": "2021-03-20T16:17:26.337000",
      "content": "<p>Thanks for all the sharing! Seems like a big architecture is common in this classification stage. I achieved 0.992 with EffNetB4 trained on 1024-&gt;512. Using B5/B6 performs worst in my case, albeit the heavy augmentations used. Somehow the 2-class classifier provided <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">in this notebook</a> (AUC 0.98) gives me the best score in public LB.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1246317,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-20T16:36:04",
          "content": "<p>yeah, for me too, i used this <a href=\"https://www.kaggle.com/mrinath/2-class-classifier-pipeline-using-effnet\" target=\"_blank\">notebook</a> for train my own 2 class filter and get auc &gt; 0.99 in my 5 folds, but it is not better than the notebook you mentioned. ☹️</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246328,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-03-20T16:49:08.290000",
          "content": "<p>You can try ensemble learning. I mixed the 2-class classifier provided in this <a href=\"https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter\" target=\"_blank\">notebook</a> with my model to get better score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1246329,
          "author_name": "Linus Johansson",
          "author_url": "",
          "post_date": "2021-03-20T16:49:45.277000",
          "content": "<p><a href=\"https://www.kaggle.com/angqx95\" target=\"_blank\">@angqx95</a> <a href=\"https://www.kaggle.com/adrielcabral\" target=\"_blank\">@adrielcabral</a> The same for me, that notebook gives me the best score. Have you guys been able find better parameters or are you using <br>\nlow_thr  = 0.08<br>\nhigh_thr = 0.95<br>\n?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1247623,
          "author_name": "adriel cabral",
          "author_url": "",
          "post_date": "2021-03-21T22:05:28.847000",
          "content": "<p><a href=\"https://www.kaggle.com/linusj79\" target=\"_blank\">@linusj79</a> sorry, i just see now your comment, i use low_thr = 0.08 high_thr = 0.95</p>\n<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> u can say how much your mixed 2 class improve lb ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1246214,
      "author_name": "Hannes Öhler",
      "author_url": "",
      "post_date": "2021-03-20T14:44:58.017000",
      "content": "<p>I also got 0.992 (efficientnet b8 with adversarial training).</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1246163,
      "author_name": "Sunghyun Jun",
      "author_url": "",
      "post_date": "2021-03-20T13:34:47.330000",
      "content": "<p>efficientnet b0 224px - auc 0.987<br>\nefficientnet b5 456px - auc 0.992<br>\nefficientnet b5 600px - auc 0.993<br>\nefficientnet b6 528px - auc 0.992<br>\nresnet200d 600px - auc 0.992 (short training, I'm trying to train longer)</p>\n<p>These are results of my experiments.<br>\nI didn't use hard augmentations like cutmix.<br>\nAlso I tried train with concat dataset of VBD and NIH.</p>\n<p>And I'm going to</p>\n<ul>\n<li>train completely concat dataset of VBD and NIH</li>\n<li>increase image size</li>\n<li>try other architectures (nfnet, seresnet, etc ?)</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 1246305,
          "author_name": "aqx",
          "author_url": "",
          "post_date": "2021-03-20T16:19:58.163000",
          "content": "<p>I didn't attempt cutmix as i felt that it could potentially retrieve a crop of an abnormal image that doesn't include any abnormality, thereby creating unnecessary noise.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1246547,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2021-03-20T20:43:36.073000",
      "content": "<p>Does the classification really help? For us, it doesn't help at all.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1246689,
          "author_name": "aqx",
          "author_url": "",
          "post_date": "2021-03-21T02:24:38.217000",
          "content": "<p>You mean the score doesn't vary with different classifiers? Or you didnt have a complete separate 2class pipeline?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1246234,
      "author_name": "Alien",
      "author_url": "",
      "post_date": "2021-03-20T15:06:57.317000",
      "content": "<p>I use Efficientnet b7(noisy-student) to get the best auc.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1248009": "Kindly remind you that the public test set account for ONLY 10% of the whole test dataset and that amount of data will 1) introduce a big shakeup with blind trust and no robust local cross-validation; 2) doesn't have enough samples to show the differences between 2 models that have 0.1% or 0.2% different in AUC, which is already a very skewed metric in opinion.",
    "1246031": "My valid AUC for classification on single fold can't seem to break the 0.992. Was wondering what are some of the architecture people are using. I felt that a good 2-class classifier is pretty important (maybe im wrong) but there isn't much discussion on the classification stage",
    "1246297": "Thanks for all the sharing! Seems like a big architecture is common in this classification stage. I achieved 0.992 with EffNetB4 trained on 1024->512. Using B5/B6 performs worst in my case, albeit the heavy augmentations used. Somehow the 2-class classifier provided [in this notebook](https://www.kaggle.com/awsaf49/vinbigdata-2-class-filter) (AUC 0.98) gives me the best score in public LB.",
    "1246214": "I also got 0.992 (efficientnet b8 with adversarial training).",
    "1246163": "efficientnet b0 224px - auc 0.987\nefficientnet b5 456px - auc 0.992\nefficientnet b5 600px - auc 0.993\nefficientnet b6 528px - auc 0.992\nresnet200d 600px - auc 0.992 (short training, I'm trying to train longer)\n\nThese are results of my experiments.\nI didn't use hard augmentations like cutmix.\nAlso I tried train with concat dataset of VBD and NIH.\n\nAnd I'm going to\n- train completely concat dataset of VBD and NIH\n- increase image size\n- try other architectures (nfnet, seresnet, etc ?)",
    "1246547": "Does the classification really help? For us, it doesn't help at all.",
    "1246234": "I use Efficientnet b7(noisy-student) to get the best auc."
  }
}