{
  "id": 183986,
  "title": "One series per study? + Question on Label Definition",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183986",
  "author_name": "Qian Cao",
  "post_date": "2020-09-18T20:01:44.487000",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello,</p>\n<ol>\n<li><p>Is it guaranteed that each study will only have one series (image volume)?</p></li>\n<li><p>For label:</p></li>\n</ol>\n<p><strong>negative_exam_for_pe</strong> - exam-level, whether there are any images in the study that have PE present.</p>\n<p>are studies containing images with PE present supposed to be labelled 1? or 0? Since the label name suggests exam \"negative\" for PE, which is a bit confusing.</p>\n<p>Thanks,</p>\n<p>Q</p>",
  "messages": [
    {
      "id": 1016303,
      "postDate": "2020-09-18T20:01:44.487Z",
      "content": "<p>Hello,</p>\n<ol>\n<li><p>Is it guaranteed that each study will only have one series (image volume)?</p></li>\n<li><p>For label:</p></li>\n</ol>\n<p><strong>negative_exam_for_pe</strong> - exam-level, whether there are any images in the study that have PE present.</p>\n<p>are studies containing images with PE present supposed to be labelled 1? or 0? Since the label name suggests exam \"negative\" for PE, which is a bit confusing.</p>\n<p>Thanks,</p>\n<p>Q</p>",
      "rawMarkdown": "Hello,\n\n1. Is it guaranteed that each study will only have one series (image volume)?\n\n2. For label:\n\n**negative_exam_for_pe** - exam-level, whether there are any images in the study that have PE present.\n\nare studies containing images with PE present supposed to be labelled 1? or 0? Since the label name suggests exam \"negative\" for PE, which is a bit confusing.\n\nThanks,\n\nQ",
      "votes": 4
    },
    {
      "id": 1016338,
      "postDate": "2020-09-18T20:59:05.157Z",
      "content": "<p>My understanding:</p>\n<p>There are three possibilities for an exam level result:</p>\n<ul>\n<li>negative</li>\n<li>indeterminate</li>\n<li>positive</li>\n</ul>\n<p>Looking at the fields:<br>\nnegative_exam_for_pe - same for every image in the exam</p>\n<ul>\n<li><p>0 (not negative, which means either indeterminate or positive)</p></li>\n<li><p>1 (negative) - there are no images with a \"pe_present_on_image\" set</p>\n<p>if \"negative_exam_for_pe\" is 0 that could be indeterminate or positive:<br>\n             - indeterminate - 0 - there is at least one image with a \"pe_present_on_image\" set<br>\n             - indeterminate - 1 - no images with \"pe_present_on_image\" set</p></li>\n</ul>\n<p>So a \"positive\" on the exam level means</p>\n<ul>\n<li>negative_exam_for_pe == 0</li>\n<li>indeterminate == 0</li>\n<li>at least one image with \"pe_present_on_image\" set</li>\n</ul>\n<p>Some basic Python code:</p>\n<p>df = pd.read_csv(mypath + '/train.csv')<br>\ndel df['pe_present_on_image']<br>\ndel df['SOPInstanceUID']<br>\ndf2 = df.drop_duplicates()<br>\ndf2['negative_exam_for_pe'].sum()<br>\ndf2['indeterminate'].sum()<br>\ndf2[df2['negative_exam_for_pe'] + df2['indeterminate'] == 0]</p>\n<p>Yields -<br>\nNegative = 4911<br>\nIndeterminate - 157<br>\nPositive - 2211</p>",
      "rawMarkdown": "My understanding:\n\nThere are three possibilities for an exam level result:\n - negative\n - indeterminate\n - positive\n\nLooking at the fields:\nnegative_exam_for_pe - same for every image in the exam\n\n   - 0 (not negative, which means either indeterminate or positive)\n   - 1 (negative) - there are no images with a \"pe_present_on_image\" set\n\n    if \"negative_exam_for_pe\" is 0 that could be indeterminate or positive:\n                 - indeterminate - 0 - there is at least one image with a \"pe_present_on_image\" set\n                 - indeterminate - 1 - no images with \"pe_present_on_image\" set\n                \nSo a \"positive\" on the exam level means\n - negative_exam_for_pe == 0\n - indeterminate == 0\n - at least one image with \"pe_present_on_image\" set\n\nSome basic Python code:\n\ndf = pd.read_csv(mypath + '/train.csv')\ndel df['pe_present_on_image']\ndel df['SOPInstanceUID']\ndf2 = df.drop_duplicates()\ndf2['negative_exam_for_pe'].sum()\ndf2['indeterminate'].sum()\ndf2[df2['negative_exam_for_pe'] + df2['indeterminate'] == 0]\n             \nYields -\nNegative = 4911\nIndeterminate - 157\nPositive - 2211\n",
      "votes": 1,
      "replies": [
        {
          "id": 1052631,
          "postDate": "2020-10-18T03:29:07.203Z",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> does order of ids matter in sub csv, i get completely opp results on lb compared to cv …</p>",
          "rawMarkdown": "@richardepstein does order of ids matter in sub csv, i get completely opp results on lb compared to cv ...",
          "replies": [
            {
              "id": 1053176,
              "postDate": "2020-10-18T16:51:19.267Z",
              "content": "<p>I do not think so. However, I always update rows in the sample_submission.csv file, so my submit order is the same as the sample_submission.csv order.</p>",
              "rawMarkdown": "I do not think so. However, I always update rows in the sample_submission.csv file, so my submit order is the same as the sample_submission.csv order."
            },
            {
              "id": 1053212,
              "postDate": "2020-10-18T17:41:53.577Z",
              "content": "<p>ok thanks <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a>  will it be right assumption to make ,if pe present on image is 1 then it is at location mentioned in CT labels  like right,central,left</p>",
              "rawMarkdown": "ok thanks @richardepstein  will it be right assumption to make ,if pe present on image is 1 then it is at location mentioned in CT labels  like right,central,left"
            },
            {
              "id": 1053273,
              "postDate": "2020-10-18T19:38:14.700Z",
              "content": "<p>right, left, central are STUDY level tags.</p>\n<p>pe_present_on_image is an IMAGE level tag.</p>\n<p>So if a STUDY has 10 images with pe_present_on_image, and has STUDY level \"right\" and \"left\", then you don't know which images are \"right\" and which are \"left\" and which are both.</p>\n<p>If you only have \"right\" for the STUDY, than any images with PE are \"right\".</p>\n<p>But you are not asked on an IMAGE level where the PE is. Only at the STUDY level.</p>",
              "rawMarkdown": "right, left, central are STUDY level tags.\n\npe_present_on_image is an IMAGE level tag.\n\nSo if a STUDY has 10 images with pe_present_on_image, and has STUDY level \"right\" and \"left\", then you don't know which images are \"right\" and which are \"left\" and which are both.\n\nIf you only have \"right\" for the STUDY, than any images with PE are \"right\".\n\nBut you are not asked on an IMAGE level where the PE is. Only at the STUDY level."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1016338,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-18T20:59:05.157000",
      "content": "<p>My understanding:</p>\n<p>There are three possibilities for an exam level result:</p>\n<ul>\n<li>negative</li>\n<li>indeterminate</li>\n<li>positive</li>\n</ul>\n<p>Looking at the fields:<br>\nnegative_exam_for_pe - same for every image in the exam</p>\n<ul>\n<li><p>0 (not negative, which means either indeterminate or positive)</p></li>\n<li><p>1 (negative) - there are no images with a \"pe_present_on_image\" set</p>\n<p>if \"negative_exam_for_pe\" is 0 that could be indeterminate or positive:<br>\n             - indeterminate - 0 - there is at least one image with a \"pe_present_on_image\" set<br>\n             - indeterminate - 1 - no images with \"pe_present_on_image\" set</p></li>\n</ul>\n<p>So a \"positive\" on the exam level means</p>\n<ul>\n<li>negative_exam_for_pe == 0</li>\n<li>indeterminate == 0</li>\n<li>at least one image with \"pe_present_on_image\" set</li>\n</ul>\n<p>Some basic Python code:</p>\n<p>df = pd.read_csv(mypath + '/train.csv')<br>\ndel df['pe_present_on_image']<br>\ndel df['SOPInstanceUID']<br>\ndf2 = df.drop_duplicates()<br>\ndf2['negative_exam_for_pe'].sum()<br>\ndf2['indeterminate'].sum()<br>\ndf2[df2['negative_exam_for_pe'] + df2['indeterminate'] == 0]</p>\n<p>Yields -<br>\nNegative = 4911<br>\nIndeterminate - 157<br>\nPositive - 2211</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1052631,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-18T03:29:07.203000",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> does order of ids matter in sub csv, i get completely opp results on lb compared to cv …</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1053176,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-18T16:51:19.267000",
              "content": "<p>I do not think so. However, I always update rows in the sample_submission.csv file, so my submit order is the same as the sample_submission.csv order.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1053212,
              "author_name": "Jaideep",
              "author_url": "",
              "post_date": "2020-10-18T17:41:53.577000",
              "content": "<p>ok thanks <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a>  will it be right assumption to make ,if pe present on image is 1 then it is at location mentioned in CT labels  like right,central,left</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 1053273,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-18T19:38:14.700000",
              "content": "<p>right, left, central are STUDY level tags.</p>\n<p>pe_present_on_image is an IMAGE level tag.</p>\n<p>So if a STUDY has 10 images with pe_present_on_image, and has STUDY level \"right\" and \"left\", then you don't know which images are \"right\" and which are \"left\" and which are both.</p>\n<p>If you only have \"right\" for the STUDY, than any images with PE are \"right\".</p>\n<p>But you are not asked on an IMAGE level where the PE is. Only at the STUDY level.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1016303": "Hello,\n\n1. Is it guaranteed that each study will only have one series (image volume)?\n\n2. For label:\n\n**negative_exam_for_pe** - exam-level, whether there are any images in the study that have PE present.\n\nare studies containing images with PE present supposed to be labelled 1? or 0? Since the label name suggests exam \"negative\" for PE, which is a bit confusing.\n\nThanks,\n\nQ",
    "1016338": "My understanding:\n\nThere are three possibilities for an exam level result:\n - negative\n - indeterminate\n - positive\n\nLooking at the fields:\nnegative_exam_for_pe - same for every image in the exam\n\n   - 0 (not negative, which means either indeterminate or positive)\n   - 1 (negative) - there are no images with a \"pe_present_on_image\" set\n\n    if \"negative_exam_for_pe\" is 0 that could be indeterminate or positive:\n                 - indeterminate - 0 - there is at least one image with a \"pe_present_on_image\" set\n                 - indeterminate - 1 - no images with \"pe_present_on_image\" set\n                \nSo a \"positive\" on the exam level means\n - negative_exam_for_pe == 0\n - indeterminate == 0\n - at least one image with \"pe_present_on_image\" set\n\nSome basic Python code:\n\ndf = pd.read_csv(mypath + '/train.csv')\ndel df['pe_present_on_image']\ndel df['SOPInstanceUID']\ndf2 = df.drop_duplicates()\ndf2['negative_exam_for_pe'].sum()\ndf2['indeterminate'].sum()\ndf2[df2['negative_exam_for_pe'] + df2['indeterminate'] == 0]\n             \nYields -\nNegative = 4911\nIndeterminate - 157\nPositive - 2211\n"
  }
}