{
  "id": 219221,
  "title": "Preferable radiologist's id in the test dataset? ",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/219221",
  "author_name": "corochann",
  "post_date": "2021-02-14T00:44:22.985000",
  "votes": 35,
  "comment_count": 11,
  "views": 0,
  "content": "<h2>Motivation</h2>\n<p>If you read the paper <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">“VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”</a> for this competition dataset, you would understand the annotation process for training data and test data is different.</p>\n<blockquote>\n  <p>For the test set, 5 radiologists involved into a two-stage<br>\n  labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other<br>\n  radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated<br>\n  with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and<br>\n  resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.</p>\n</blockquote>\n<p>Ground truth data in the test dataset is finalized by 2 reviewers, which is not available in the training dataset. Therefore we need to <strong>estimate which annotations is more preferable by these 2 reviewers</strong>.</p>\n<p>I wonder there is specific radiologists whose annotation tend to be accepted or rejected.</p>\n<h2>Data check and experimental setting</h2>\n<p>In the training dataset, 4394 abnormal images out of 15,000 images.<br>\n<strong>And actually 4146 abnormal images (94.3%) are annotated by R8, R9 and R10 radiologists (these 3 radiologists always annotate together).</strong></p>\n<p>I experimented by removing R8, R9 or R10's annotation, to see the impact of LB score to understand there's preferable annotation by <code>rad_id</code>.<br>\nIn the following, I experimented based on the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.</p>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Exp setting</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Use all annotations</td>\n<td>0.232</td>\n</tr>\n<tr>\n<td>Remove \"R8\" annotation</td>\n<td>0.218</td>\n</tr>\n<tr>\n<td>Remove \"R9\" annotation</td>\n<td>0.225</td>\n</tr>\n<tr>\n<td>Remove \"R10\" annotation</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>Remove \"R8\"&amp; \"R9\" annotation</td>\n<td>0.197</td>\n</tr>\n<tr>\n<td>Remove \"R9\"&amp; \"R10\" annotation</td>\n<td>0.217</td>\n</tr>\n<tr>\n<td>Remove \"R8\"&amp; \"R10\" annotation</td>\n<td>0.224</td>\n</tr>\n</tbody>\n</table>\n<p>Based on this experiment, <strong>using all the annotation performed the best</strong>.<br>\nEspecially, removing R8 or R9 annotation worse the LB score.</p>\n<h2>Observation</h2>\n<p>It seems the importance of the annotation is \"R8\", \"R9\", \"R10\".<br>\nI guess test dataset is annotated selecting specific bbox, instead of creating new ensembled bbox. My idea is to applying NMS with the confidence score weighted with the order \"R8\", \"R9\", \"R10\"  improves the score.</p>\n<h2>Code</h2>\n<p>You can experiment it by changing the <code>get_vinbigdata_dicts</code> method by following from the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.<br>\nFor example, <code>ignore_rad_ids = [\"R8\", \"R9\"]</code> is to Remove \"R8\"&amp; \"R9\" annotation.</p>\n<pre><code>def get_vinbigdata_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    train_data_type: str = \"original\",\n    use_cache: bool = True,\n    debug: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    use_class14: bool = False,\n    ignore_rad_ids: Optional[List[str]] = None\n):\n    debug_str = f\"_debug{int(debug)}\"\n    train_data_type_str = f\"_{train_data_type}\"\n    class14_str = f\"_14class{int(use_class14)}\"\n    ignore_rad_ids_str = \"\" if ignore_rad_ids is None else \"_ig\" + \"\".join(ignore_rad_ids)\n    cache_path = imgdir / f\"dataset_dicts_cache{train_data_type_str}{class14_str}{ignore_rad_ids_str}{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        train_meta = pd.read_csv(imgdir / \"train_meta.csv\")\n        if debug:\n            train_meta = train_meta.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train_meta.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"train\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        n_ignored = 0\n        for index, train_meta_row in tqdm(train_meta.iterrows(), total=len(train_meta)):\n            record = {}\n\n            image_id, height, width = train_meta_row.values\n            filename = str(imgdir / \"train\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"image_id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"class_id\"]\n                rad_id = row[\"rad_id\"]\n                if class_id == 14:\n                    # It is \"No finding\"\n                    if use_class14:\n                        # Use this No finding class with the bbox covering all image area.\n                        bbox_resized = [0, 0, resized_width, resized_height]\n                        obj = {\n                            \"bbox\": bbox_resized,\n                            \"bbox_mode\": BoxMode.XYXY_ABS,\n                            \"category_id\": class_id,\n                            \"rad_id\": rad_id,\n                        }\n                        objs.append(obj)\n                    else:\n                        # This annotator does not find anything, skip.\n                        pass\n                    # If \"No finding\", all other annotation is also \"No finding\".\n                    # Skip counting duplicated \"No finding\".\n                    break\n                else:\n                    if ignore_rad_ids is not None and rad_id in ignore_rad_ids:\n                        # Skip this rad_id's annotation...\n                        n_ignored += 1\n                        continue\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                        \"rad_id\": rad_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n        print(f\"ignore_rad_ids {ignore_rad_ids}, # ignored: {n_ignored}\")\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n</code></pre>",
  "messages": [
    {
      "id": 1199588,
      "postDate": "2021-02-14T00:44:22.987Z",
      "content": "<h2>Motivation</h2>\n<p>If you read the paper <a href=\"https://arxiv.org/pdf/2012.15029.pdf\" target=\"_blank\">“VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”</a> for this competition dataset, you would understand the annotation process for training data and test data is different.</p>\n<blockquote>\n  <p>For the test set, 5 radiologists involved into a two-stage<br>\n  labeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other<br>\n  radiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated<br>\n  with each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and<br>\n  resolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.</p>\n</blockquote>\n<p>Ground truth data in the test dataset is finalized by 2 reviewers, which is not available in the training dataset. Therefore we need to <strong>estimate which annotations is more preferable by these 2 reviewers</strong>.</p>\n<p>I wonder there is specific radiologists whose annotation tend to be accepted or rejected.</p>\n<h2>Data check and experimental setting</h2>\n<p>In the training dataset, 4394 abnormal images out of 15,000 images.<br>\n<strong>And actually 4146 abnormal images (94.3%) are annotated by R8, R9 and R10 radiologists (these 3 radiologists always annotate together).</strong></p>\n<p>I experimented by removing R8, R9 or R10's annotation, to see the impact of LB score to understand there's preferable annotation by <code>rad_id</code>.<br>\nIn the following, I experimented based on the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.</p>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>Exp setting</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Use all annotations</td>\n<td>0.232</td>\n</tr>\n<tr>\n<td>Remove \"R8\" annotation</td>\n<td>0.218</td>\n</tr>\n<tr>\n<td>Remove \"R9\" annotation</td>\n<td>0.225</td>\n</tr>\n<tr>\n<td>Remove \"R10\" annotation</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>Remove \"R8\"&amp; \"R9\" annotation</td>\n<td>0.197</td>\n</tr>\n<tr>\n<td>Remove \"R9\"&amp; \"R10\" annotation</td>\n<td>0.217</td>\n</tr>\n<tr>\n<td>Remove \"R8\"&amp; \"R10\" annotation</td>\n<td>0.224</td>\n</tr>\n</tbody>\n</table>\n<p>Based on this experiment, <strong>using all the annotation performed the best</strong>.<br>\nEspecially, removing R8 or R9 annotation worse the LB score.</p>\n<h2>Observation</h2>\n<p>It seems the importance of the annotation is \"R8\", \"R9\", \"R10\".<br>\nI guess test dataset is annotated selecting specific bbox, instead of creating new ensembled bbox. My idea is to applying NMS with the confidence score weighted with the order \"R8\", \"R9\", \"R10\"  improves the score.</p>\n<h2>Code</h2>\n<p>You can experiment it by changing the <code>get_vinbigdata_dicts</code> method by following from the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.<br>\nFor example, <code>ignore_rad_ids = [\"R8\", \"R9\"]</code> is to Remove \"R8\"&amp; \"R9\" annotation.</p>\n<pre><code>def get_vinbigdata_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    train_data_type: str = \"original\",\n    use_cache: bool = True,\n    debug: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    use_class14: bool = False,\n    ignore_rad_ids: Optional[List[str]] = None\n):\n    debug_str = f\"_debug{int(debug)}\"\n    train_data_type_str = f\"_{train_data_type}\"\n    class14_str = f\"_14class{int(use_class14)}\"\n    ignore_rad_ids_str = \"\" if ignore_rad_ids is None else \"_ig\" + \"\".join(ignore_rad_ids)\n    cache_path = imgdir / f\"dataset_dicts_cache{train_data_type_str}{class14_str}{ignore_rad_ids_str}{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        train_meta = pd.read_csv(imgdir / \"train_meta.csv\")\n        if debug:\n            train_meta = train_meta.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train_meta.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"train\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        n_ignored = 0\n        for index, train_meta_row in tqdm(train_meta.iterrows(), total=len(train_meta)):\n            record = {}\n\n            image_id, height, width = train_meta_row.values\n            filename = str(imgdir / \"train\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"image_id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"class_id\"]\n                rad_id = row[\"rad_id\"]\n                if class_id == 14:\n                    # It is \"No finding\"\n                    if use_class14:\n                        # Use this No finding class with the bbox covering all image area.\n                        bbox_resized = [0, 0, resized_width, resized_height]\n                        obj = {\n                            \"bbox\": bbox_resized,\n                            \"bbox_mode\": BoxMode.XYXY_ABS,\n                            \"category_id\": class_id,\n                            \"rad_id\": rad_id,\n                        }\n                        objs.append(obj)\n                    else:\n                        # This annotator does not find anything, skip.\n                        pass\n                    # If \"No finding\", all other annotation is also \"No finding\".\n                    # Skip counting duplicated \"No finding\".\n                    break\n                else:\n                    if ignore_rad_ids is not None and rad_id in ignore_rad_ids:\n                        # Skip this rad_id's annotation...\n                        n_ignored += 1\n                        continue\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                        \"rad_id\": rad_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n        print(f\"ignore_rad_ids {ignore_rad_ids}, # ignored: {n_ignored}\")\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n</code></pre>",
      "rawMarkdown": "## Motivation\n\nIf you read the paper [“VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”](https://arxiv.org/pdf/2012.15029.pdf) for this competition dataset, you would understand the annotation process for training data and test data is different.\n\n> For the test set, 5 radiologists involved into a two-stage\nlabeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other\nradiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated\nwith each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and\nresolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\n\nGround truth data in the test dataset is finalized by 2 reviewers, which is not available in the training dataset. Therefore we need to **estimate which annotations is more preferable by these 2 reviewers**.\n\nI wonder there is specific radiologists whose annotation tend to be accepted or rejected.\n\n## Data check and experimental setting\nIn the training dataset, 4394 abnormal images out of 15,000 images.\n**And actually 4146 abnormal images (94.3%) are annotated by R8, R9 and R10 radiologists (these 3 radiologists always annotate together).**\n\nI experimented by removing R8, R9 or R10's annotation, to see the impact of LB score to understand there's preferable annotation by `rad_id`.\nIn the following, I experimented based on the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train).\n\n## Results\n| Exp setting | LB score |\n| --- | --- |\n| Use all annotations | 0.232 |\n| Remove \"R8\" annotation | 0.218 |\n| Remove \"R9\" annotation | 0.225 |\n| Remove \"R10\" annotation | 0.231 |\n| Remove \"R8\"& \"R9\" annotation | 0.197 |\n| Remove \"R9\"& \"R10\" annotation | 0.217 |\n| Remove \"R8\"& \"R10\" annotation | 0.224 |\n\nBased on this experiment, **using all the annotation performed the best**.\nEspecially, removing R8 or R9 annotation worse the LB score.\n\n## Observation\n\nIt seems the importance of the annotation is \"R8\", \"R9\", \"R10\".\nI guess test dataset is annotated selecting specific bbox, instead of creating new ensembled bbox. My idea is to applying NMS with the confidence score weighted with the order \"R8\", \"R9\", \"R10\"  improves the score.\n\n## Code\n\nYou can experiment it by changing the `get_vinbigdata_dicts` method by following from the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train).\nFor example, `ignore_rad_ids = [\"R8\", \"R9\"]` is to Remove \"R8\"& \"R9\" annotation.\n\n```\ndef get_vinbigdata_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    train_data_type: str = \"original\",\n    use_cache: bool = True,\n    debug: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    use_class14: bool = False,\n    ignore_rad_ids: Optional[List[str]] = None\n):\n    debug_str = f\"_debug{int(debug)}\"\n    train_data_type_str = f\"_{train_data_type}\"\n    class14_str = f\"_14class{int(use_class14)}\"\n    ignore_rad_ids_str = \"\" if ignore_rad_ids is None else \"_ig\" + \"\".join(ignore_rad_ids)\n    cache_path = imgdir / f\"dataset_dicts_cache{train_data_type_str}{class14_str}{ignore_rad_ids_str}{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        train_meta = pd.read_csv(imgdir / \"train_meta.csv\")\n        if debug:\n            train_meta = train_meta.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train_meta.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"train\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        n_ignored = 0\n        for index, train_meta_row in tqdm(train_meta.iterrows(), total=len(train_meta)):\n            record = {}\n\n            image_id, height, width = train_meta_row.values\n            filename = str(imgdir / \"train\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"image_id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"class_id\"]\n                rad_id = row[\"rad_id\"]\n                if class_id == 14:\n                    # It is \"No finding\"\n                    if use_class14:\n                        # Use this No finding class with the bbox covering all image area.\n                        bbox_resized = [0, 0, resized_width, resized_height]\n                        obj = {\n                            \"bbox\": bbox_resized,\n                            \"bbox_mode\": BoxMode.XYXY_ABS,\n                            \"category_id\": class_id,\n                            \"rad_id\": rad_id,\n                        }\n                        objs.append(obj)\n                    else:\n                        # This annotator does not find anything, skip.\n                        pass\n                    # If \"No finding\", all other annotation is also \"No finding\".\n                    # Skip counting duplicated \"No finding\".\n                    break\n                else:\n                    if ignore_rad_ids is not None and rad_id in ignore_rad_ids:\n                        # Skip this rad_id's annotation...\n                        n_ignored += 1\n                        continue\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                        \"rad_id\": rad_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n        print(f\"ignore_rad_ids {ignore_rad_ids}, # ignored: {n_ignored}\")\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n```",
      "votes": 35
    },
    {
      "id": 1239974,
      "postDate": "2021-03-16T07:02:46.147Z",
      "content": "<p>i would think that:</p>\n<ol>\n<li>since that we are evaluating at iou&lt;0.4, size of the box is less important, except for small box</li>\n<li>we need to detect all objects in the image. you need to find a way to reduce multiple training annotations to single using \"some method\". assume we try different methods:<ul>\n<li>choice 1: x1,x2,y1,y2, score  (using method 1)</li>\n<li>choice 2: x1,x2,y1,y2, score  (using method 2)<br>\n…</li>\n<li>choice N: x1,x2,y1,y2, score  (using method 3)</li></ul></li>\n</ol>\n<p>you just need to find one that correlates best to the LB score.<br>\ne.g. method1 can just be average, method2 can take top-K, method3 can be accept everything, etc…</p>",
      "rawMarkdown": "i would think that:\n1.  since that we are evaluating at iou<0.4, size of the box is less important, except for small box\n2. we need to detect all objects in the image. you need to find a way to reduce multiple training annotations to single using \"some method\". assume we try different methods:\n   - choice 1: x1,x2,y1,y2, score  (using method 1)\n   - choice 2: x1,x2,y1,y2, score  (using method 2)\n...\n   - choice N: x1,x2,y1,y2, score  (using method 3)\n\n\nyou just need to find one that correlates best to the LB score.\ne.g. method1 can just be average, method2 can take top-K, method3 can be accept everything, etc...",
      "votes": 3,
      "replies": [
        {
          "id": 1240566,
          "postDate": "2021-03-16T13:48:42.573Z",
          "content": "<p>That is a very rational proposal. But, with only 300 images in public LB, it's very hard to confirm any significant CV-&gt;LB correlation to confirm the benefit of any merging method without overfitting public LB. Any LB change can possibly be related to the very low number of rare  annotations in the randomly selected 300 images.</p>\n<p>When I bootstrap a local val set of 3000 images in N groups of 300 images, I always get large variations of mAP locally.</p>\n<p>In summary, we have different distributions between training and LB sets presumably mostly related to ground truth labels but we have minimal signal to understand the distribution shift.</p>",
          "rawMarkdown": "That is a very rational proposal. But, with only 300 images in public LB, it's very hard to confirm any significant CV->LB correlation to confirm the benefit of any merging method without overfitting public LB. Any LB change can possibly be related to the very low number of rare  annotations in the randomly selected 300 images.\n\nWhen I bootstrap a local val set of 3000 images in N groups of 300 images, I always get large variations of mAP locally.\n\nIn summary, we have different distributions between training and LB sets presumably mostly related to ground truth labels but we have minimal signal to understand the distribution shift.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1211056,
      "postDate": "2021-02-19T23:33:42.023Z",
      "content": "<p>As a radiologist myself, it is pretty interesting to see how badly we need a non-human second reader,<br>\npossibly free of bias… and also to see how hard it is going to be to achieve it  =)</p>",
      "rawMarkdown": "As a radiologist myself, it is pretty interesting to see how badly we need a non-human second reader,\npossibly free of bias... and also to see how hard it is going to be to achieve it  =)",
      "votes": 3
    },
    {
      "id": 1199952,
      "postDate": "2021-02-14T09:21:42.523Z",
      "content": "<p>The while thing gets only more complicated by some of the reviewers <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181\" target=\"_blank\">only ever reviewing images that were believed to have no findings</a> up front.</p>\n<p>Personally, I would not expect a single radiologist to be preferable to the whole group, unless they are much more skilled than the others. It's a bit like enseembling models: it nearly always helps as long as the models are similar enough in performance (and ideally make very different mistakes - or just randomly make mistakes, which to some extent probably happens with radiologists here).</p>",
      "rawMarkdown": "The while thing gets only more complicated by some of the reviewers [only ever reviewing images that were believed to have no findings](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181) up front.\n\nPersonally, I would not expect a single radiologist to be preferable to the whole group, unless they are much more skilled than the others. It's a bit like enseembling models: it nearly always helps as long as the models are similar enough in performance (and ideally make very different mistakes - or just randomly make mistakes, which to some extent probably happens with radiologists here).",
      "votes": 4
    },
    {
      "id": 1199876,
      "postDate": "2021-02-14T07:27:17.887Z",
      "content": "<p>Nice analysis, and thank you for sharing the result.</p>",
      "rawMarkdown": "Nice analysis, and thank you for sharing the result.",
      "votes": 1
    },
    {
      "id": 1201496,
      "postDate": "2021-02-15T12:41:40.213Z",
      "content": "<p>Interesting analysis, thank you for sharing!<br>\nBut this is largely an estimation since we only have 300 images on the public LB…unfortunately the same experiment cannot be conducted on a local validation set due to labeling method difference</p>",
      "rawMarkdown": "Interesting analysis, thank you for sharing!\nBut this is largely an estimation since we only have 300 images on the public LB...unfortunately the same experiment cannot be conducted on a local validation set due to labeling method difference",
      "votes": 2,
      "replies": [
        {
          "id": 1203131,
          "postDate": "2021-02-15T21:33:00.800Z",
          "content": "<p>That's right!</p>",
          "rawMarkdown": "That's right!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1237432,
      "postDate": "2021-03-14T06:21:02.513Z",
      "content": "<p>what about the result if just use the \"R8 R9 R10\" annotations?</p>",
      "rawMarkdown": "what about the result if just use the \"R8 R9 R10\" annotations?",
      "replies": [
        {
          "id": 1239626,
          "postDate": "2021-03-15T21:25:07.200Z",
          "content": "<p>what's the purpose? some images annotation are done by other radiologists, you will simply lose ground truth of these images.</p>",
          "rawMarkdown": "what's the purpose? some images annotation are done by other radiologists, you will simply lose ground truth of these images."
        }
      ]
    },
    {
      "id": 1205665,
      "postDate": "2021-02-16T21:31:31.313Z",
      "content": "<p>Does it improve the score?</p>",
      "rawMarkdown": "Does it improve the score?",
      "replies": [
        {
          "id": 1205694,
          "postDate": "2021-02-16T22:13:21.500Z",
          "content": "<p>You can read \"Results\" section, so far there is no improvement. But you can get an idea what is important from this experiment.</p>",
          "rawMarkdown": "You can read \"Results\" section, so far there is no improvement. But you can get an idea what is important from this experiment."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1239974,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-03-16T07:02:46.147000",
      "content": "<p>i would think that:</p>\n<ol>\n<li>since that we are evaluating at iou&lt;0.4, size of the box is less important, except for small box</li>\n<li>we need to detect all objects in the image. you need to find a way to reduce multiple training annotations to single using \"some method\". assume we try different methods:<ul>\n<li>choice 1: x1,x2,y1,y2, score  (using method 1)</li>\n<li>choice 2: x1,x2,y1,y2, score  (using method 2)<br>\n…</li>\n<li>choice N: x1,x2,y1,y2, score  (using method 3)</li></ul></li>\n</ol>\n<p>you just need to find one that correlates best to the LB score.<br>\ne.g. method1 can just be average, method2 can take top-K, method3 can be accept everything, etc…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1240566,
          "author_name": "Alexandre Cadrin-Chênevert",
          "author_url": "",
          "post_date": "2021-03-16T13:48:42.573000",
          "content": "<p>That is a very rational proposal. But, with only 300 images in public LB, it's very hard to confirm any significant CV-&gt;LB correlation to confirm the benefit of any merging method without overfitting public LB. Any LB change can possibly be related to the very low number of rare  annotations in the randomly selected 300 images.</p>\n<p>When I bootstrap a local val set of 3000 images in N groups of 300 images, I always get large variations of mAP locally.</p>\n<p>In summary, we have different distributions between training and LB sets presumably mostly related to ground truth labels but we have minimal signal to understand the distribution shift.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1211056,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2021-02-19T23:33:42.023000",
      "content": "<p>As a radiologist myself, it is pretty interesting to see how badly we need a non-human second reader,<br>\npossibly free of bias… and also to see how hard it is going to be to achieve it  =)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1199952,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-14T09:21:42.523000",
      "content": "<p>The while thing gets only more complicated by some of the reviewers <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181\" target=\"_blank\">only ever reviewing images that were believed to have no findings</a> up front.</p>\n<p>Personally, I would not expect a single radiologist to be preferable to the whole group, unless they are much more skilled than the others. It's a bit like enseembling models: it nearly always helps as long as the models are similar enough in performance (and ideally make very different mistakes - or just randomly make mistakes, which to some extent probably happens with radiologists here).</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1199876,
      "author_name": "JIN",
      "author_url": "",
      "post_date": "2021-02-14T07:27:17.887000",
      "content": "<p>Nice analysis, and thank you for sharing the result.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1201496,
      "author_name": "InDSweTrust",
      "author_url": "",
      "post_date": "2021-02-15T12:41:40.213000",
      "content": "<p>Interesting analysis, thank you for sharing!<br>\nBut this is largely an estimation since we only have 300 images on the public LB…unfortunately the same experiment cannot be conducted on a local validation set due to labeling method difference</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1203131,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-02-15T21:33:00.800000",
          "content": "<p>That's right!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1237432,
      "author_name": "qyyyyyf",
      "author_url": "",
      "post_date": "2021-03-14T06:21:02.513000",
      "content": "<p>what about the result if just use the \"R8 R9 R10\" annotations?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1239626,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-03-15T21:25:07.200000",
          "content": "<p>what's the purpose? some images annotation are done by other radiologists, you will simply lose ground truth of these images.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1205665,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2021-02-16T21:31:31.313000",
      "content": "<p>Does it improve the score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1205694,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-02-16T22:13:21.500000",
          "content": "<p>You can read \"Results\" section, so far there is no improvement. But you can get an idea what is important from this experiment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1199588": "## Motivation\n\nIf you read the paper [“VinDr-CXR: An open dataset of chest X-rays with radiologist's annotations”](https://arxiv.org/pdf/2012.15029.pdf) for this competition dataset, you would understand the annotation process for training data and test data is different.\n\n> For the test set, 5 radiologists involved into a two-stage\nlabeling process. During the first stage, each image was independently annotated by 3 radiologists. In the second stage, 2 other\nradiologists, who have a higher level of experience, reviewed the annotations of the 3 previous annotators and communicated\nwith each other in order to decide the final labels. The disagreements among initial annotators were carefully discussed and\nresolved by the 2 reviewers. Finally, the consensus of their opinions will serve as reference ground-truth.\n\nGround truth data in the test dataset is finalized by 2 reviewers, which is not available in the training dataset. Therefore we need to **estimate which annotations is more preferable by these 2 reviewers**.\n\nI wonder there is specific radiologists whose annotation tend to be accepted or rejected.\n\n## Data check and experimental setting\nIn the training dataset, 4394 abnormal images out of 15,000 images.\n**And actually 4146 abnormal images (94.3%) are annotated by R8, R9 and R10 radiologists (these 3 radiologists always annotate together).**\n\nI experimented by removing R8, R9 or R10's annotation, to see the impact of LB score to understand there's preferable annotation by `rad_id`.\nIn the following, I experimented based on the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train).\n\n## Results\n| Exp setting | LB score |\n| --- | --- |\n| Use all annotations | 0.232 |\n| Remove \"R8\" annotation | 0.218 |\n| Remove \"R9\" annotation | 0.225 |\n| Remove \"R10\" annotation | 0.231 |\n| Remove \"R8\"& \"R9\" annotation | 0.197 |\n| Remove \"R9\"& \"R10\" annotation | 0.217 |\n| Remove \"R8\"& \"R10\" annotation | 0.224 |\n\nBased on this experiment, **using all the annotation performed the best**.\nEspecially, removing R8 or R9 annotation worse the LB score.\n\n## Observation\n\nIt seems the importance of the annotation is \"R8\", \"R9\", \"R10\".\nI guess test dataset is annotated selecting specific bbox, instead of creating new ensembled bbox. My idea is to applying NMS with the confidence score weighted with the order \"R8\", \"R9\", \"R10\"  improves the score.\n\n## Code\n\nYou can experiment it by changing the `get_vinbigdata_dicts` method by following from the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train).\nFor example, `ignore_rad_ids = [\"R8\", \"R9\"]` is to Remove \"R8\"& \"R9\" annotation.\n\n```\ndef get_vinbigdata_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    train_data_type: str = \"original\",\n    use_cache: bool = True,\n    debug: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    use_class14: bool = False,\n    ignore_rad_ids: Optional[List[str]] = None\n):\n    debug_str = f\"_debug{int(debug)}\"\n    train_data_type_str = f\"_{train_data_type}\"\n    class14_str = f\"_14class{int(use_class14)}\"\n    ignore_rad_ids_str = \"\" if ignore_rad_ids is None else \"_ig\" + \"\".join(ignore_rad_ids)\n    cache_path = imgdir / f\"dataset_dicts_cache{train_data_type_str}{class14_str}{ignore_rad_ids_str}{debug_str}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        train_meta = pd.read_csv(imgdir / \"train_meta.csv\")\n        if debug:\n            train_meta = train_meta.iloc[:500]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train_meta.loc[0, \"image_id\"]\n        image_path = str(imgdir / \"train\" / f\"{image_id}.png\")\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        n_ignored = 0\n        for index, train_meta_row in tqdm(train_meta.iterrows(), total=len(train_meta)):\n            record = {}\n\n            image_id, height, width = train_meta_row.values\n            filename = str(imgdir / \"train\" / f\"{image_id}.png\")\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = resized_height\n            record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"image_id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"class_id\"]\n                rad_id = row[\"rad_id\"]\n                if class_id == 14:\n                    # It is \"No finding\"\n                    if use_class14:\n                        # Use this No finding class with the bbox covering all image area.\n                        bbox_resized = [0, 0, resized_width, resized_height]\n                        obj = {\n                            \"bbox\": bbox_resized,\n                            \"bbox_mode\": BoxMode.XYXY_ABS,\n                            \"category_id\": class_id,\n                            \"rad_id\": rad_id,\n                        }\n                        objs.append(obj)\n                    else:\n                        # This annotator does not find anything, skip.\n                        pass\n                    # If \"No finding\", all other annotation is also \"No finding\".\n                    # Skip counting duplicated \"No finding\".\n                    break\n                else:\n                    if ignore_rad_ids is not None and rad_id in ignore_rad_ids:\n                        # Skip this rad_id's annotation...\n                        n_ignored += 1\n                        continue\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                        \"rad_id\": rad_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n        print(f\"ignore_rad_ids {ignore_rad_ids}, # ignored: {n_ignored}\")\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n```",
    "1239974": "i would think that:\n1.  since that we are evaluating at iou<0.4, size of the box is less important, except for small box\n2. we need to detect all objects in the image. you need to find a way to reduce multiple training annotations to single using \"some method\". assume we try different methods:\n   - choice 1: x1,x2,y1,y2, score  (using method 1)\n   - choice 2: x1,x2,y1,y2, score  (using method 2)\n...\n   - choice N: x1,x2,y1,y2, score  (using method 3)\n\n\nyou just need to find one that correlates best to the LB score.\ne.g. method1 can just be average, method2 can take top-K, method3 can be accept everything, etc...",
    "1211056": "As a radiologist myself, it is pretty interesting to see how badly we need a non-human second reader,\npossibly free of bias... and also to see how hard it is going to be to achieve it  =)",
    "1199952": "The while thing gets only more complicated by some of the reviewers [only ever reviewing images that were believed to have no findings](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/216181) up front.\n\nPersonally, I would not expect a single radiologist to be preferable to the whole group, unless they are much more skilled than the others. It's a bit like enseembling models: it nearly always helps as long as the models are similar enough in performance (and ideally make very different mistakes - or just randomly make mistakes, which to some extent probably happens with radiologists here).",
    "1199876": "Nice analysis, and thank you for sharing the result.",
    "1201496": "Interesting analysis, thank you for sharing!\nBut this is largely an estimation since we only have 300 images on the public LB...unfortunately the same experiment cannot be conducted on a local validation set due to labeling method difference",
    "1237432": "what about the result if just use the \"R8 R9 R10\" annotations?",
    "1205665": "Does it improve the score?"
  }
}