{
  "id": 226669,
  "title": "The ap on test data is inconsistent with that on validation subset",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/226669",
  "author_name": "dmyfighting",
  "post_date": "2021-03-17T08:59:05.119000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I split the training data into train/validation subsets. The ap on test data is inconsistent with that on validation subset. Does someone else have the same question?</p>",
  "messages": [
    {
      "id": 1241877,
      "postDate": "2021-03-17T08:59:05.120Z",
      "content": "<p>I split the training data into train/validation subsets. The ap on test data is inconsistent with that on validation subset. Does someone else have the same question?</p>",
      "rawMarkdown": "I split the training data into train/validation subsets. The ap on test data is inconsistent with that on validation subset. Does someone else have the same question?",
      "votes": 5
    },
    {
      "id": 1242242,
      "postDate": "2021-03-17T13:47:33.197Z",
      "content": "<p>I believe almost everyone are experiencing a significant shift in performance between CV and LB. If we assume that the images in train and test sets are coming from the same randomly assigned generating process, then the shift is likely explained by 2 potential reasons related to labels:</p>\n<p>1) Difference in ground truth definition. Consensus of 3+2 rads in test set; Customizable merge of 3 rads in train test. This challenge is all about finding a good solution to this shift.</p>\n<p>2) Difference in \"No findings\" prevalence between your val set and the test set. This can be corrected easily if your val set contains only images with positive findings.</p>",
      "rawMarkdown": "I believe almost everyone are experiencing a significant shift in performance between CV and LB. If we assume that the images in train and test sets are coming from the same randomly assigned generating process, then the shift is likely explained by 2 potential reasons related to labels:\n\n1) Difference in ground truth definition. Consensus of 3+2 rads in test set; Customizable merge of 3 rads in train test. This challenge is all about finding a good solution to this shift.\n\n2) Difference in \"No findings\" prevalence between your val set and the test set. This can be corrected easily if your val set contains only images with positive findings.\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3243847,
          "postDate": "2025-07-07T15:18:06.700Z",
          "content": "<p>Hi Alexandre</p>\n<p>I know this was posted long ago but just wanted to ask your opinion for current research using this dataset.</p>\n<p>In Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:<br>\n<a href=\"https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\" target=\"_blank\">https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf</a></p>\n<p>Here it seems that also report results from an internal validation set of 3,000 images in Table 1:<br>\n<a href=\"https://arxiv.org/pdf/2410.21969\" target=\"_blank\">https://arxiv.org/pdf/2410.21969</a></p>\n<p>It is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.</p>\n<p>I was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.</p>\n<p>What do you think?</p>",
          "rawMarkdown": "Hi Alexandre\n\nI know this was posted long ago but just wanted to ask your opinion for current research using this dataset.\n\nIn Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:\nhttps://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\n\nHere it seems that also report results from an internal validation set of 3,000 images in Table 1:\nhttps://arxiv.org/pdf/2410.21969\n\nIt is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.\n\nI was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.\n\nWhat do you think?"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1242242,
      "author_name": "Alexandre Cadrin-Chênevert",
      "author_url": "",
      "post_date": "2021-03-17T13:47:33.197000",
      "content": "<p>I believe almost everyone are experiencing a significant shift in performance between CV and LB. If we assume that the images in train and test sets are coming from the same randomly assigned generating process, then the shift is likely explained by 2 potential reasons related to labels:</p>\n<p>1) Difference in ground truth definition. Consensus of 3+2 rads in test set; Customizable merge of 3 rads in train test. This challenge is all about finding a good solution to this shift.</p>\n<p>2) Difference in \"No findings\" prevalence between your val set and the test set. This can be corrected easily if your val set contains only images with positive findings.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3243847,
          "author_name": "Joshua Bruton",
          "author_url": "",
          "post_date": "2025-07-07T15:18:06.700000",
          "content": "<p>Hi Alexandre</p>\n<p>I know this was posted long ago but just wanted to ask your opinion for current research using this dataset.</p>\n<p>In Table 1 of this paper you can see they use an internal holdout of 3000 images from the 15,000 images to report performance:<br>\n<a href=\"https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf\" target=\"_blank\">https://www.medrxiv.org/content/10.1101/2021.09.28.21264286v1.full.pdf</a></p>\n<p>Here it seems that also report results from an internal validation set of 3,000 images in Table 1:<br>\n<a href=\"https://arxiv.org/pdf/2410.21969\" target=\"_blank\">https://arxiv.org/pdf/2410.21969</a></p>\n<p>It is strange to me that the literature is refraining from using the Test dataset as defined by VinDr simply because it makes the results look bad. The datasets feel as though they are from completely different distributions when I use them.</p>\n<p>I was thinking of either doing like the above two papers and just ignoring the test set, or creating a new stratified split from the full 18,000 with perfectly balanced class representation in train/val/test.</p>\n<p>What do you think?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1241877": "I split the training data into train/validation subsets. The ap on test data is inconsistent with that on validation subset. Does someone else have the same question?",
    "1242242": "I believe almost everyone are experiencing a significant shift in performance between CV and LB. If we assume that the images in train and test sets are coming from the same randomly assigned generating process, then the shift is likely explained by 2 potential reasons related to labels:\n\n1) Difference in ground truth definition. Consensus of 3+2 rads in test set; Customizable merge of 3 rads in train test. This challenge is all about finding a good solution to this shift.\n\n2) Difference in \"No findings\" prevalence between your val set and the test set. This can be corrected easily if your val set contains only images with positive findings.\n\n\n"
  }
}