{
  "id": 183473,
  "title": "Label Consistency Requirement Details",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/183473",
  "author_name": "John Mongan",
  "post_date": "2020-09-16T20:12:10.334000",
  "votes": 35,
  "comment_count": 55,
  "views": 0,
  "content": "<p>The data page for this competition describes a requirement for logical consistency of predicted labels, and includes a diagram describing the relationship between labels. </p>\n<p>For the purpose of enforcing this rule we consider a label to be “predicted” if that label is assigned a probability of &gt; 0.5. We require that the “predicted” labels be logically consistent. Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs. However, due to the complexity of the metric and the difficulty in anticipating all corner cases, we have put this rule in place to reduce the ability to game the metric with nonsensical outputs.</p>\n<p>Some specific examples of how this rule plays out:</p>\n<p>At the image level, any image with predicted probability &gt;  0.5 is considered as being positive for PE will count as a positive image</p>\n<p>At the exam level, we have</p>\n<ol>\n<li>Negative, Indeterminate, (Positive) and it can only be one of these. If any image is predicted positive, there cannot also be a predicted probability of Negative &gt; 0.5 nor can there be a predicted probability of Indeterminate &gt;  0.5.</li>\n</ol>\n<p>Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5</p>\n<ol>\n<li><p>Right, left, central -- if any image is predicted positive (p &gt; 0.5) then at least one of these labels must be assigned p &gt; 0.5; more than one of these labels may be assigned p &gt; 0.5. When no images are predicted positive, then none of these labels may be assigned p &gt; 0.5</p></li>\n<li><p>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive. </p></li>\n</ol>\n<ul>\n<li>if any image on the exam is positive, one of these must have p &gt;  0.5</li>\n<li>both cannot have p &gt;  0.5</li>\n</ul>\n<ol>\n<li>Acute, Chronic, Acute &amp; Chronic -- it cannot be both chronic &amp; acute and chronic so </li>\n</ol>\n<ul>\n<li>only one can have p &gt;  0.5</li>\n<li>it is also possible that neither has p &gt; 0.5  </li>\n<li>in other words, it is inconsistent to say chronic has p &gt; 0.5 and acute &amp; chronic has p &gt;  0.5.</li>\n</ul>\n<p><strong>EDIT:</strong> The code that will be used to check compliance with these requirements is available in this <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">notebook</a>.</p>",
  "messages": [
    {
      "id": 1013615,
      "postDate": "2020-09-16T20:12:10.333Z",
      "content": "<p>The data page for this competition describes a requirement for logical consistency of predicted labels, and includes a diagram describing the relationship between labels. </p>\n<p>For the purpose of enforcing this rule we consider a label to be “predicted” if that label is assigned a probability of &gt; 0.5. We require that the “predicted” labels be logically consistent. Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs. However, due to the complexity of the metric and the difficulty in anticipating all corner cases, we have put this rule in place to reduce the ability to game the metric with nonsensical outputs.</p>\n<p>Some specific examples of how this rule plays out:</p>\n<p>At the image level, any image with predicted probability &gt;  0.5 is considered as being positive for PE will count as a positive image</p>\n<p>At the exam level, we have</p>\n<ol>\n<li>Negative, Indeterminate, (Positive) and it can only be one of these. If any image is predicted positive, there cannot also be a predicted probability of Negative &gt; 0.5 nor can there be a predicted probability of Indeterminate &gt;  0.5.</li>\n</ol>\n<p>Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5</p>\n<ol>\n<li><p>Right, left, central -- if any image is predicted positive (p &gt; 0.5) then at least one of these labels must be assigned p &gt; 0.5; more than one of these labels may be assigned p &gt; 0.5. When no images are predicted positive, then none of these labels may be assigned p &gt; 0.5</p></li>\n<li><p>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive. </p></li>\n</ol>\n<ul>\n<li>if any image on the exam is positive, one of these must have p &gt;  0.5</li>\n<li>both cannot have p &gt;  0.5</li>\n</ul>\n<ol>\n<li>Acute, Chronic, Acute &amp; Chronic -- it cannot be both chronic &amp; acute and chronic so </li>\n</ol>\n<ul>\n<li>only one can have p &gt;  0.5</li>\n<li>it is also possible that neither has p &gt; 0.5  </li>\n<li>in other words, it is inconsistent to say chronic has p &gt; 0.5 and acute &amp; chronic has p &gt;  0.5.</li>\n</ul>\n<p><strong>EDIT:</strong> The code that will be used to check compliance with these requirements is available in this <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "The data page for this competition describes a requirement for logical consistency of predicted labels, and includes a diagram describing the relationship between labels. \n\nFor the purpose of enforcing this rule we consider a label to be “predicted” if that label is assigned a probability of > 0.5. We require that the “predicted” labels be logically consistent. Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs. However, due to the complexity of the metric and the difficulty in anticipating all corner cases, we have put this rule in place to reduce the ability to game the metric with nonsensical outputs.\n\nSome specific examples of how this rule plays out:\n\nAt the image level, any image with predicted probability >  0.5 is considered as being positive for PE will count as a positive image\n\nAt the exam level, we have\n1. Negative, Indeterminate, (Positive) and it can only be one of these. If any image is predicted positive, there cannot also be a predicted probability of Negative > 0.5 nor can there be a predicted probability of Indeterminate >  0.5.\n\nSimilarly, if no image is positive (p > 0.5), then there must be one and only one negative or indeterminate with p > 0.5\n\n2. Right, left, central -- if any image is predicted positive (p > 0.5) then at least one of these labels must be assigned p > 0.5; more than one of these labels may be assigned p > 0.5. When no images are predicted positive, then none of these labels may be assigned p > 0.5\n\n3. RV/LV ratio. It can be only one of these and it must be present if at least one image is positive. \n- if any image on the exam is positive, one of these must have p >  0.5\n- both cannot have p >  0.5\n\n4. Acute, Chronic, Acute & Chronic -- it cannot be both chronic & acute and chronic so \n- only one can have p >  0.5\n- it is also possible that neither has p > 0.5  \n- in other words, it is inconsistent to say chronic has p > 0.5 and acute & chronic has p >  0.5.\n\n**EDIT:** The code that will be used to check compliance with these requirements is available in this [notebook](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check).",
      "votes": 34
    },
    {
      "id": 1013821,
      "postDate": "2020-09-17T02:23:07.643Z",
      "content": "<p>To avoid a fiasco where one of the winners ends up getting disqualified because of a single prediction that ends up violating these requirements, it should be built into the submission system to throw an error if these requirements are not met. </p>\n<p>That would be the kind thing to do. </p>",
      "rawMarkdown": "To avoid a fiasco where one of the winners ends up getting disqualified because of a single prediction that ends up violating these requirements, it should be built into the submission system to throw an error if these requirements are not met. \n\nThat would be the kind thing to do. \n\n",
      "votes": 13,
      "replies": [
        {
          "id": 1013955,
          "postDate": "2020-09-17T05:05:28.860Z",
          "content": "<p>Agreed. <br>\nThere is always a chance of misunderstanding/not noticing these requirements. Would be very nice if it's implemented at kaggle side.</p>",
          "rawMarkdown": "Agreed. \nThere is always a chance of misunderstanding/not noticing these requirements. Would be very nice if it's implemented at kaggle side."
        },
        {
          "id": 1014729,
          "postDate": "2020-09-17T16:53:27.993Z",
          "content": "<p>Automated checking of all submissions would be ideal, but unfortunately wasn't logistically feasible for this competition. We will be checking the winning submissions for consistency. The goal is not to trip people up; we don't want to disqualify anyone. At the same time, given the complexity of the metric, if logically inconsistent outputs are discovered that are scored by the metric, we don't want to be in the position of awarding algorithms that produce nonsensical output as winners. We will be working to make sure that the script that will be used to validate consistency of the winners is available to all contestants during the competition so they can check their own outputs for consistency (see <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> notebook later in this thread). We recognize this is less convenient than automated checking, but we hope this will reduce the chances of unhappy surprises.</p>",
          "rawMarkdown": "Automated checking of all submissions would be ideal, but unfortunately wasn't logistically feasible for this competition. We will be checking the winning submissions for consistency. The goal is not to trip people up; we don't want to disqualify anyone. At the same time, given the complexity of the metric, if logically inconsistent outputs are discovered that are scored by the metric, we don't want to be in the position of awarding algorithms that produce nonsensical output as winners. We will be working to make sure that the script that will be used to validate consistency of the winners is available to all contestants during the competition so they can check their own outputs for consistency (see @kozodoi notebook later in this thread). We recognize this is less convenient than automated checking, but we hope this will reduce the chances of unhappy surprises.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1014525,
      "postDate": "2020-09-17T14:03:35.490Z",
      "content": "<p>These consistency requirements, while problematic from the modeling viewpoint, are important from the perspective of \"Explainable AI\". </p>\n<p>Many \"models\" give impressive statistical results but are very difficult to move into actual clinical practice.</p>\n<p>For a model to be useful, it must be medically consistent, otherwise the physicians/health care providers will not trust it.</p>\n<p>For instance, if you report 8 images with 99% probably of PE, but your overall score for the exam is 10%, what is the clinician to think? What is their next action. Treat? Ignore?</p>\n<p>Do I treat a 30% probably of PE? At 50%? 80%?</p>\n<p>A prediction of 40% Positive, 30% Negative and 30% Indeterminate gives no guidance to treatment.</p>\n<p>We have all seen heat maps where the model predicts \"Pneumonia\" and the heat map is on the Right/Left marker.</p>\n<p>This competition is not asking for bounding boxes, but that is essentially what Right/Left/Central are. It doesn't help to say \"Positive for Pulmonary Embolism\" but not be able to show the clinician where it is.</p>\n<p>RV/LV ratio is essentially a whole separate Competition hidden within this one. RV/LV ratio can be elevated in the setting of pulmonary embolism and can signify the need for more intensive monitoring and treatment. It really is a binary classifier because by definition either the RV/LV is greater or less than one. The challenge for the model is first you have to find the heart and then you have to determine the RV/LV ratio. Most images might contribute nothing because they don't include the heart.</p>\n<p>While we have some work to do, \"Explainability\" and medical consistency are crucial features to get these model accepted and used.</p>",
      "rawMarkdown": "These consistency requirements, while problematic from the modeling viewpoint, are important from the perspective of \"Explainable AI\". \n\nMany \"models\" give impressive statistical results but are very difficult to move into actual clinical practice.\n\nFor a model to be useful, it must be medically consistent, otherwise the physicians/health care providers will not trust it.\n\nFor instance, if you report 8 images with 99% probably of PE, but your overall score for the exam is 10%, what is the clinician to think? What is their next action. Treat? Ignore?\n\nDo I treat a 30% probably of PE? At 50%? 80%?\n\nA prediction of 40% Positive, 30% Negative and 30% Indeterminate gives no guidance to treatment.\n\nWe have all seen heat maps where the model predicts \"Pneumonia\" and the heat map is on the Right/Left marker.\n\nThis competition is not asking for bounding boxes, but that is essentially what Right/Left/Central are. It doesn't help to say \"Positive for Pulmonary Embolism\" but not be able to show the clinician where it is.\n\nRV/LV ratio is essentially a whole separate Competition hidden within this one. RV/LV ratio can be elevated in the setting of pulmonary embolism and can signify the need for more intensive monitoring and treatment. It really is a binary classifier because by definition either the RV/LV is greater or less than one. The challenge for the model is first you have to find the heart and then you have to determine the RV/LV ratio. Most images might contribute nothing because they don't include the heart.\n\nWhile we have some work to do, \"Explainability\" and medical consistency are crucial features to get these model accepted and used.",
      "votes": 10
    },
    {
      "id": 1016032,
      "postDate": "2020-09-18T15:59:46.777Z",
      "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> you stated above <code>Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs.</code> But as the metric give 0 weight to the ‘PE Present on Image’ of images from series with no PE. Hence the \"punishment\" of the metric for predicting 'Negative for PE =1' and ‘PE Present on Image=1’ for images in studies with no PE is 0! </p>",
      "rawMarkdown": "@anthracene you stated above `Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs.` But as the metric give 0 weight to the ‘PE Present on Image’ of images from series with no PE. Hence the \"punishment\" of the metric for predicting 'Negative for PE =1' and ‘PE Present on Image=1’ for images in studies with no PE is 0! ",
      "votes": 6
    },
    {
      "id": 1024457,
      "postDate": "2020-09-23T22:05:30.893Z",
      "content": "<p>Please take a look, as the host has added an officially-vetted version of <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">the code</a> that will be used to check compliance with these requirements, appended to the end of this original post.</p>",
      "rawMarkdown": "Please take a look, as the host has added an officially-vetted version of [the code](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check) that will be used to check compliance with these requirements, appended to the end of this original post.",
      "votes": 4,
      "replies": [
        {
          "id": 1035686,
          "postDate": "2020-10-03T00:33:52.007Z",
          "content": "<p>In the code, the rule, <code>If any image is predicted positive, there cannot also be a predicted probability of Negative &gt; 0.5 nor can there be a predicted probability of Indeterminate &gt; 0.5.</code>,  is not checked., doesn't it?</p>",
          "rawMarkdown": "In the code, the rule, ```If any image is predicted positive, there cannot also be a predicted probability of Negative > 0.5 nor can there be a predicted probability of Indeterminate > 0.5.```,  is not checked., doesn't it?",
          "votes": 2,
          "replies": [
            {
              "id": 1035697,
              "postDate": "2020-10-03T01:03:47.167Z",
              "content": "<p>Our understanding is that right now, none of the consistency rules are checked or enforced on the leaderboard. As discussed above, they plan on checking the top ten on the Private Leaderboard at the end of the competition. Others have posted Notebooks that implement the described checks.</p>",
              "rawMarkdown": "Our understanding is that right now, none of the consistency rules are checked or enforced on the leaderboard. As discussed above, they plan on checking the top ten on the Private Leaderboard at the end of the competition. Others have posted Notebooks that implement the described checks.",
              "votes": 2
            }
          ]
        },
        {
          "id": 1035850,
          "postDate": "2020-10-03T06:40:05.873Z",
          "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> looks like you are correct this condition isn't check in the released code. <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> could you please take a look.</p>",
          "rawMarkdown": "@osciiart looks like you are correct this condition isn't check in the released code. @anthracene @kozodoi could you please take a look."
        },
        {
          "id": 1035861,
          "postDate": "2020-10-03T06:59:43.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> you are right, this rule was missing. I implemented it in <a href=\"https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements?scriptVersionId=43935750\" target=\"_blank\">Version 5</a> of my consistency check notebook. </p>\n<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> if it looks fine to you, maybe you could update the official version as well.</p>",
          "rawMarkdown": "@osciiart @yuval6967 you are right, this rule was missing. I implemented it in [Version 5](https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements?scriptVersionId=43935750) of my consistency check notebook. \n\n@anthracene if it looks fine to you, maybe you could update the official version as well.",
          "votes": 2
        },
        {
          "id": 1041777,
          "postDate": "2020-10-07T23:03:37.793Z",
          "content": "<p>Thank you; we've updated the official version to add enforcement of this rule that was missing.</p>",
          "rawMarkdown": "Thank you; we've updated the official version to add enforcement of this rule that was missing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1013864,
      "postDate": "2020-09-17T03:13:40.027Z",
      "content": "<blockquote>\n  <p>Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5<br>\n  When no images are predicted positive, then none of these labels may be assigned p &gt; 0.5</p>\n</blockquote>\n<p>I am confused.  positive prob &lt; 0.5 for all images does not mean that exam level negative prob &gt; 0.5.</p>\n<p>Data desc page says that:</p>\n<blockquote>\n  <p>negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.</p>\n</blockquote>\n<p>So   <code>[Prob of Exam level Negative] = Π [prob that i-th image is nagative] = Π(1 - P_positive_i)</code>.<br>\nThen <code>P_positive_i &lt; 0.5 for all i-th image</code> does not mean <code>[Prob of Exam level Negative] &gt; 0.5</code><br>\nFor example, <code>P_positive = 0.4 for all i and number of image in exam is 2</code>, then <code>[Prob of Exam level Negative] = 0.6*0.6=0.36 &lt; 0.5</code></p>\n<p>Could you give me some explanation?</p>\n<p>[Update]<br>\nI also think below is not valid requirement.</p>\n<blockquote>\n  <p>Right, left, central -- if any image is predicted positive (p &gt; 0.5) then at least one of these labels must be assigned p &gt; 0.5</p>\n</blockquote>\n<p>Forecasting tomorrow's weather as \"30% shiny, 30% snowy, 40% rainy\" is definitely valid prediction. If you want at least one positive out of 3 candidates, you can use threshold of 1/3 instead of 1/2. There exists no reason to use 1/2.  <br>\nIt is ok to require <code>P(left)+P(central)+P(right) &gt;= 1</code> and use threthold of 1/3 to get at least one postive.  But host requires more</p>",
      "rawMarkdown": "> Similarly, if no image is positive (p > 0.5), then there must be one and only one negative or indeterminate with p > 0.5\n> When no images are predicted positive, then none of these labels may be assigned p > 0.5\n\nI am confused.  positive prob < 0.5 for all images does not mean that exam level negative prob > 0.5.\n\nData desc page says that:\n> negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.\n\nSo   `[Prob of Exam level Negative] = Π [prob that i-th image is nagative] = Π(1 - P_positive_i)`.\nThen `P_positive_i < 0.5 for all i-th image` does not mean `[Prob of Exam level Negative] > 0.5`\nFor example, `P_positive = 0.4 for all i and number of image in exam is 2`, then `[Prob of Exam level Negative] = 0.6*0.6=0.36 < 0.5`\n\nCould you give me some explanation?\n\n\n[Update]\nI also think below is not valid requirement.\n> Right, left, central -- if any image is predicted positive (p > 0.5) then at least one of these labels must be assigned p > 0.5\n\nForecasting tomorrow's weather as \"30% shiny, 30% snowy, 40% rainy\" is definitely valid prediction. If you want at least one positive out of 3 candidates, you can use threshold of 1/3 instead of 1/2. There exists no reason to use 1/2.  \nIt is ok to require `P(left)+P(central)+P(right) >= 1` and use threthold of 1/3 to get at least one postive.  But host requires more",
      "votes": 1,
      "replies": [
        {
          "id": 1014743,
          "postDate": "2020-09-17T17:01:18.960Z",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> 's post in this thread does an excellent job of explaining the rationale for these rules</p>",
          "rawMarkdown": "@richardepstein 's post in this thread does an excellent job of explaining the rationale for these rules"
        },
        {
          "id": 1015480,
          "postDate": "2020-09-18T08:06:33.863Z",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> Thanks for your resnponce.</p>\n<p>What host wants can be summarized as \"at least one positive\".<br>\nThat can be easily implemented in the application-phase. The apps does not even have to show probabilities to users.<br>\nI believe that modelling-phase should not have a responsibility for it.</p>\n<p>Especially, logloss metrics implicitly assumes calibrated probability as a prediction. Statistically meaningless requirements distort the evaluation.<br>\nI think that the requirement should be amended to make this competition better.</p>\n<p>I admit (some) users may feel nervous to \"30% left, 30% central, 40% right\" prediction.<br>\nBut \"at least 50% prob requirements\" does not resolve the issue and it is not related to \"logical consistency\". There is no difference between 30%, 33.3% and 50%.</p>",
          "rawMarkdown": "@anthracene Thanks for your resnponce.\n\nWhat host wants can be summarized as \"at least one positive\".\nThat can be easily implemented in the application-phase. The apps does not even have to show probabilities to users.\nI believe that modelling-phase should not have a responsibility for it.\n\nEspecially, logloss metrics implicitly assumes calibrated probability as a prediction. Statistically meaningless requirements distort the evaluation.\nI think that the requirement should be amended to make this competition better.\n\nI admit (some) users may feel nervous to \"30% left, 30% central, 40% right\" prediction.\nBut \"at least 50% prob requirements\" does not resolve the issue and it is not related to \"logical consistency\". There is no difference between 30%, 33.3% and 50%.",
          "votes": 2
        },
        {
          "id": 1050925,
          "postDate": "2020-10-15T23:03:14.890Z",
          "content": "<p>[Prob of Exam level Negative] =/= Π [prob that i-th image is nagative] !!!!<br>\nΠ [prob that i-th image is nagative] = is the probability that <em>all</em> images are negative for PE.<br>\na better approximation for [Prob of Exam level Negative] is max probability that  i-th image is nagative.</p>\n<p>Moreover, there can be more than one PE in a single lung. Therefore, Right, left, central are NOT mutually exclusive (e.g. there can be a right AND a central PE in the same lung).</p>",
          "rawMarkdown": "[Prob of Exam level Negative] =/= Π [prob that i-th image is nagative] !!!!\nΠ [prob that i-th image is nagative] = is the probability that *all* images are negative for PE.\na better approximation for [Prob of Exam level Negative] is max probability that  i-th image is nagative.\n\nMoreover, there can be more than one PE in a single lung. Therefore, Right, left, central are NOT mutually exclusive (e.g. there can be a right AND a central PE in the same lung).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1058230,
      "postDate": "2020-10-23T13:31:24.130Z",
      "content": "<p>Good, but what about the training set ? There are such inconsistencies in it too, for example cases when neither of <code>pe_present_on_image, negative_exam_for_pe, indeterminate</code> are equal to one. Perhaps this has already been mentioned in discussions or notebooks, and I missed it. If it is the case, I apologize.</p>\n<p>I used this code : </p>\n<pre><code>labels = ['pe_present_on_image', 'negative_exam_for_pe', 'indeterminate']\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n</code></pre>\n<p>We obtain : </p>\n<p>Total nb of images : 1790594<br>\nImages with ONE of the three categories mentioned above, equal to one : 1344365<br>\nImages with NONE of the three categories mentioned above, equal to one : 446229</p>",
      "rawMarkdown": "Good, but what about the training set ? There are such inconsistencies in it too, for example cases when neither of `pe_present_on_image, negative_exam_for_pe, indeterminate` are equal to one. Perhaps this has already been mentioned in discussions or notebooks, and I missed it. If it is the case, I apologize.\n\nI used this code : \n\n```\nlabels = ['pe_present_on_image', 'negative_exam_for_pe', 'indeterminate']\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n```\n\nWe obtain : \n\nTotal nb of images : 1790594\nImages with ONE of the three categories mentioned above, equal to one : 1344365\nImages with NONE of the three categories mentioned above, equal to one : 446229",
      "replies": [
        {
          "id": 1058277,
          "postDate": "2020-10-23T14:25:49.810Z",
          "content": "<p>You need to do this. For each study/exam, you need to calculate <code>max_pe_present</code> in order to know whether that study is positive or not.</p>\n<pre><code>labels = ['max_pe_present', 'negative_exam_for_pe', 'indeterminate']\ndf_train['max_pe_present'] = \\   \n  df_train.groupby('StudyInstanceUID').pe_present_on_image.transform('max')\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n</code></pre>\n<p>Then it prints 1790594 and 1790594</p>",
          "rawMarkdown": "You need to do this. For each study/exam, you need to calculate `max_pe_present` in order to know whether that study is positive or not.\n\n    labels = ['max_pe_present', 'negative_exam_for_pe', 'indeterminate']\n    df_train['max_pe_present'] = \\   \n      df_train.groupby('StudyInstanceUID').pe_present_on_image.transform('max')\n    df_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\n    print(df_train.shape[0])\n    df_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\n    print(df_train_1.shape[0])\n\nThen it prints 1790594 and 1790594"
        },
        {
          "id": 1058285,
          "postDate": "2020-10-23T14:33:11.613Z",
          "content": "<p>Indeed, you are right. Thank you. Of course, we need to have at least one image per study with   pe_present_on_image=1 in order to have a positive study.</p>",
          "rawMarkdown": "Indeed, you are right. Thank you. Of course, we need to have at least one image per study with   pe_present_on_image=1 in order to have a positive study.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1054463,
      "postDate": "2020-10-20T00:40:43.077Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>The rules in this post (and the competition rules <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/data\" target=\"_blank\">here</a>) do not agree with the rules in the official notebook <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">here</a>. Could you please clarify which is correct?</p>\n<p>Specifically, this thread allows us to predict both <code>pe_present_on_image&lt;=0.5</code> and <code>rv_lv_ratio_gte_1 &gt; 0.5</code> for an entire exam/study whereas the notebook does not allow this.</p>\n<p>Can a patient have a right ventricle that is larger than a left ventricle and still not have pulmonary embolism?</p>",
      "rawMarkdown": "Hello @anthracene @philculliton @juliaelliott \n\nThe rules in this post (and the competition rules [here][2]) do not agree with the rules in the official notebook [here][1]. Could you please clarify which is correct?\n\nSpecifically, this thread allows us to predict both `pe_present_on_image<=0.5` and `rv_lv_ratio_gte_1 > 0.5` for an entire exam/study whereas the notebook does not allow this.\n\nCan a patient have a right ventricle that is larger than a left ventricle and still not have pulmonary embolism?\n\n[1]: https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\n[2]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/data",
      "replies": [
        {
          "id": 1055349,
          "postDate": "2020-10-20T17:26:28.260Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for raising! The host is reviewing your feedback and will respond directly as soon as possible.</p>",
          "rawMarkdown": "@cdeotte Thanks for raising! The host is reviewing your feedback and will respond directly as soon as possible.",
          "votes": 2
        },
        {
          "id": 1057644,
          "postDate": "2020-10-22T20:21:10.243Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , the host responded that if exam has all <code>pe_present_on_image&lt;=0.5</code> then <code>rv_lv_ratio_gte_1</code> and <code>rv_lv_ratio_lt_1</code> must be <code>&lt;=0.5</code>.</p>\n<p>I have another question. Which teams' submissions will be checked for Label Consistency? And what if one submission conforms but the second submission does not conform?</p>\n<p>I notice the <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/rules\" target=\"_blank\">rules</a> say below but I was wondering if you could be more specific like (1) only cash prizes will be checked (2) or all medal positions will be checked, etc.</p>\n<blockquote>\n  <p>Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page. This includes:</p>\n</blockquote>",
          "rawMarkdown": "Thanks @juliaelliott , the host responded that if exam has all `pe_present_on_image<=0.5` then `rv_lv_ratio_gte_1` and `rv_lv_ratio_lt_1` must be `<=0.5`.\n\nI have another question. Which teams' submissions will be checked for Label Consistency? And what if one submission conforms but the second submission does not conform?\n\nI notice the [rules][1] say below but I was wondering if you could be more specific like (1) only cash prizes will be checked (2) or all medal positions will be checked, etc.\n\n> Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page. This includes:\n\n[1]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/rules"
        },
        {
          "id": 1057647,
          "postDate": "2020-10-22T20:22:35.970Z",
          "content": "<p>Thanks for your response in my other <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696\" target=\"_blank\">deleted post</a>. You said</p>\n<blockquote>\n  <p><strong>Prize winners'</strong> winning submissions will be checked and disqualified/removed if non-conforming. They should not expect to rely on any secondary selected submission as a fall-back.</p>\n</blockquote>",
          "rawMarkdown": "Thanks for your response in my other [deleted post][1]. You said\n> **Prize winners'** winning submissions will be checked and disqualified/removed if non-conforming. They should not expect to rely on any secondary selected submission as a fall-back.\n\n[1]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696"
        },
        {
          "id": 1057725,
          "postDate": "2020-10-22T23:35:52.860Z",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nI'd like to know too.</p>",
          "rawMarkdown": "@juliaelliott \nI'd like to know too."
        },
        {
          "id": 1057728,
          "postDate": "2020-10-22T23:43:41.697Z",
          "content": "<p>Both questions have been answered. The host answered about <code>rv_lv_ratio_gte_1</code> below. And Julia answered about leaderboard removal <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696\" target=\"_blank\">here</a>. (Originally i asked the question in a separate topic but moved it here. I deleted it before i noticed that Julia responded).</p>",
          "rawMarkdown": "Both questions have been answered. The host answered about `rv_lv_ratio_gte_1` below. And Julia answered about leaderboard removal [here][1]. (Originally i asked the question in a separate topic but moved it here. I deleted it before i noticed that Julia responded).\n\n[1]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696"
        },
        {
          "id": 1057755,
          "postDate": "2020-10-23T00:56:22.210Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks 😄</p>",
          "rawMarkdown": "@cdeotte Thanks 😄"
        },
        {
          "id": 1057756,
          "postDate": "2020-10-23T00:57:36.197Z",
          "content": "<p>oh sry I missed it.</p>",
          "rawMarkdown": "oh sry I missed it."
        }
      ]
    },
    {
      "id": 1052550,
      "postDate": "2020-10-17T22:16:19.553Z",
      "content": "<p>From the explanation and from the notebook which implements these errors' checking, it transpires that for an exam with at least one positive image, it is allowed to have both Acute &lt;= 0.5 and Acute and Chronic &lt;= 0.5. </p>\n<p>On the other hand, in the description of the competition (the schema) it is written that for positive exam we need to have at least one of Acute or Acute and Chronic &gt; 0.5.</p>\n<p>Which statement is true ? Can someone explain ? </p>",
      "rawMarkdown": "From the explanation and from the notebook which implements these errors' checking, it transpires that for an exam with at least one positive image, it is allowed to have both Acute <= 0.5 and Acute and Chronic <= 0.5. \n\nOn the other hand, in the description of the competition (the schema) it is written that for positive exam we need to have at least one of Acute or Acute and Chronic > 0.5.\n\nWhich statement is true ? Can someone explain ? ",
      "replies": [
        {
          "id": 1052564,
          "postDate": "2020-10-17T23:31:17.297Z",
          "content": "<p>The schema below only uses the words \"at least one\" for \"right-sided\", \"left-sided\", and \"central\". Regarding \"RV/LV\" targets and \"Chronic Acute\" targets. It says \"only one\" which means \"at most one which can be zero\".</p>\n<p>To answer your answer, if you have a positive exam, you are allowed to leave both \"chronic\" and \"chronic+acute\" &lt;=0.5 for that exam. (This represents the case of just \"acute\" for that exam)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F69f7671bbaa31bbda88224e4d16a509f%2Fscheme.png?generation=1602977187941229&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "The schema below only uses the words \"at least one\" for \"right-sided\", \"left-sided\", and \"central\". Regarding \"RV/LV\" targets and \"Chronic Acute\" targets. It says \"only one\" which means \"at most one which can be zero\".\n\nTo answer your answer, if you have a positive exam, you are allowed to leave both \"chronic\" and \"chronic+acute\" <=0.5 for that exam. (This represents the case of just \"acute\" for that exam)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F69f7671bbaa31bbda88224e4d16a509f%2Fscheme.png?generation=1602977187941229&alt=media)",
          "votes": 1
        },
        {
          "id": 1052567,
          "postDate": "2020-10-17T23:35:43.283Z",
          "content": "<p>Thank you! The truth is it is a little ambiguous in the schema. It is clear now. And simplifies things.</p>",
          "rawMarkdown": "Thank you! The truth is it is a little ambiguous in the schema. It is clear now. And simplifies things.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1052538,
      "postDate": "2020-10-17T21:28:45.770Z",
      "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/jeffrudie\" target=\"_blank\">@jeffrudie</a> </p>\n<p>John, does your linked notebook (latest version 2) enforce the rules (what you've written above) correctly? The notebook says that we cannot predict <code>rv_lv_ratio_lt_1 &gt;0.5</code> nor <code>rv_lv_ratio_gte_1 &gt;0.5</code> if all <code>pe_present_on_image&lt;=0.5</code> for a given study with the following code</p>\n<pre><code>rule2b = df_neg.loc[(df_neg.rv_lv_ratio_lt_1     &gt; 0.5) | \n                (df_neg.rv_lv_ratio_gte_1    &gt; 0.5) |\n                (df_neg.central_pe           &gt; 0.5) | \n                (df_neg.rightsided_pe        &gt; 0.5) | \n                (df_neg.leftsided_pe         &gt; 0.5) |\n                (df_neg.acute_and_chronic_pe &gt; 0.5) | \n                (df_neg.chronic_pe           &gt; 0.5)].reset_index(drop = True)\n</code></pre>\n<p>I don't think this agrees with what you wrote above</p>\n<blockquote>\n  <p>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive.<br>\n  if any image on the exam is positive, one of these must have p &gt; 0.5<br>\n  both cannot have p &gt; 0.5</p>\n</blockquote>\n<p>The statement you wrote allows us to predict <code>rv_lv_ratio_lt_1 &gt;0.5</code> or <code>rv_lv_ratio_gte_1 &gt;0.5</code> if all <code>pe_present_on_image&lt;=0.5</code> for a given study. </p>\n<p>Am i correct that your code is different that your text, and if so, which rule is correct?</p>",
      "rawMarkdown": "@anthracene @jeffrudie \n\nJohn, does your linked notebook (latest version 2) enforce the rules (what you've written above) correctly? The notebook says that we cannot predict `rv_lv_ratio_lt_1 >0.5` nor `rv_lv_ratio_gte_1 >0.5` if all `pe_present_on_image<=0.5` for a given study with the following code\n  \n    rule2b = df_neg.loc[(df_neg.rv_lv_ratio_lt_1     > 0.5) | \n                    (df_neg.rv_lv_ratio_gte_1    > 0.5) |\n                    (df_neg.central_pe           > 0.5) | \n                    (df_neg.rightsided_pe        > 0.5) | \n                    (df_neg.leftsided_pe         > 0.5) |\n                    (df_neg.acute_and_chronic_pe > 0.5) | \n                    (df_neg.chronic_pe           > 0.5)].reset_index(drop = True)\n\nI don't think this agrees with what you wrote above\n\n>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive.\nif any image on the exam is positive, one of these must have p > 0.5\nboth cannot have p > 0.5\n\nThe statement you wrote allows us to predict `rv_lv_ratio_lt_1 >0.5` or `rv_lv_ratio_gte_1 >0.5` if all `pe_present_on_image<=0.5` for a given study. \n\nAm i correct that your code is different that your text, and if so, which rule is correct?\n\n[1]: https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check",
      "replies": [
        {
          "id": 1052540,
          "postDate": "2020-10-17T21:34:33.263Z",
          "content": "<p>IMHO, you need to change your notebook to the following if you want to agree with what you wrote above</p>\n<pre><code>    rule2b = df_neg.loc[((df_neg.rv_lv_ratio_lt_1     &gt; 0.5) &amp; \n                    (df_neg.rv_lv_ratio_gte_1    &gt; 0.5)) |\n                    (df_neg.central_pe           &gt; 0.5) | \n                    (df_neg.rightsided_pe        &gt; 0.5) | \n                    (df_neg.leftsided_pe         &gt; 0.5) |\n                    (df_neg.acute_and_chronic_pe &gt; 0.5) | \n                    (df_neg.chronic_pe           &gt; 0.5)].reset_index(drop = True)\n</code></pre>",
          "rawMarkdown": "IMHO, you need to change your notebook to the following if you want to agree with what you wrote above\n\n        rule2b = df_neg.loc[((df_neg.rv_lv_ratio_lt_1     > 0.5) & \n                        (df_neg.rv_lv_ratio_gte_1    > 0.5)) |\n                        (df_neg.central_pe           > 0.5) | \n                        (df_neg.rightsided_pe        > 0.5) | \n                        (df_neg.leftsided_pe         > 0.5) |\n                        (df_neg.acute_and_chronic_pe > 0.5) | \n                        (df_neg.chronic_pe           > 0.5)].reset_index(drop = True)"
        },
        {
          "id": 1057551,
          "postDate": "2020-10-22T18:43:25.160Z",
          "content": "<p>I believe the code in the notebook is correct; my English description of the rule may not be completely clear. If the study is negative, then neither rv_lv_ratio_lt_1 nor rv_lv_ratio_gte_1 may be predicted as &gt;0.5.</p>",
          "rawMarkdown": "I believe the code in the notebook is correct; my English description of the rule may not be completely clear. If the study is negative, then neither rv_lv_ratio_lt_1 nor rv_lv_ratio_gte_1 may be predicted as >0.5.",
          "votes": 1
        },
        {
          "id": 1057576,
          "postDate": "2020-10-22T19:11:10.620Z",
          "content": "<p>Ok, thanks for the update</p>",
          "rawMarkdown": "Ok, thanks for the update"
        }
      ]
    },
    {
      "id": 1049766,
      "postDate": "2020-10-14T18:58:58.657Z",
      "content": "<p>Excellent work!</p>\n<p>btw:</p>\n<ul>\n<li>v3: wrapped consistency checks into function</li>\n</ul>\n<p>should be:<br>\ncheck_consistency()</p>",
      "rawMarkdown": "Excellent work!\n\nbtw:\n\n* v3: wrapped consistency checks into~~ check_consitency()~~ function\n\nshould be:\ncheck_consistency()"
    },
    {
      "id": 1047698,
      "postDate": "2020-10-12T21:22:02.147Z",
      "content": "<p>Proud to get this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3616012%2F2689841674e17fa9bf6a21c1f041a783%2FRSNA_no_errors.png?generation=1602537705968718&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Proud to get this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3616012%2F2689841674e17fa9bf6a21c1f041a783%2FRSNA_no_errors.png?generation=1602537705968718&alt=media)",
      "replies": [
        {
          "id": 1047977,
          "postDate": "2020-10-13T05:20:45.760Z",
          "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> what about private test data?</p>",
          "rawMarkdown": "@catadanna what about private test data?"
        },
        {
          "id": 1048091,
          "postDate": "2020-10-13T07:20:28.060Z",
          "content": "<p>What do you mean by that? This is the result obtained on one of my submissions.</p>",
          "rawMarkdown": "What do you mean by that? This is the result obtained on one of my submissions."
        },
        {
          "id": 1048181,
          "postDate": "2020-10-13T09:00:36.040Z",
          "content": "<p>The reply is on the public test. You can't see the reply for the private test.</p>\n<p>What I do for the private test is something like:</p>\n<pre><code>If len(errors)==0:\n    save submission file\n</code></pre>",
          "rawMarkdown": "The reply is on the public test. You can't see the reply for the private test.\n\nWhat I do for the private test is something like:\n```\nIf len(errors)==0:\n    save submission file\n```",
          "votes": 3
        },
        {
          "id": 1048196,
          "postDate": "2020-10-13T09:17:56.190Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a>: Indeed, that is what I thought too! </p>",
          "rawMarkdown": "@yuval6967: Indeed, that is what I thought too! "
        }
      ]
    },
    {
      "id": 1024412,
      "postDate": "2020-09-23T20:53:02.867Z",
      "content": "<p>Are these rules currently enforced in the submission grader? I had submission scoring errors submitting the sample_submission (all probabilities at 0.5) and I wonder if this is the reason?</p>",
      "rawMarkdown": "Are these rules currently enforced in the submission grader? I had submission scoring errors submitting the sample_submission (all probabilities at 0.5) and I wonder if this is the reason?",
      "replies": [
        {
          "id": 1024452,
          "postDate": "2020-09-23T21:58:40.050Z",
          "content": "<p>No, unfortunately they aren't. We will manually run the top 10 private leaderboard entries against this code at the conclusion of the contest to verify compliance.</p>",
          "rawMarkdown": "No, unfortunately they aren't. We will manually run the top 10 private leaderboard entries against this code at the conclusion of the contest to verify compliance.",
          "votes": 2
        },
        {
          "id": 1029788,
          "postDate": "2020-09-28T06:13:52.423Z",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> Sorry if this has been asked before or due to my misunderstanding of the competition rule. May I ask what will happen to the top-10 if they fail the compliance check? (No offense to the current top-10!)</p>",
          "rawMarkdown": "@anthracene Sorry if this has been asked before or due to my misunderstanding of the competition rule. May I ask what will happen to the top-10 if they fail the compliance check? (No offense to the current top-10!)"
        },
        {
          "id": 1029834,
          "postDate": "2020-09-28T07:14:09.147Z",
          "content": "<p>Quoting the competition rules:</p>\n<blockquote>\n  <p>Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page.</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>A disqualified participant will be removed from the Competition leaderboard, at the Competition Sponsor and Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.</p>\n</blockquote>",
          "rawMarkdown": "Quoting the competition rules:\n\n> Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page.\n\nand\n\n> A disqualified participant will be removed from the Competition leaderboard, at the Competition Sponsor and Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.",
          "votes": 1
        },
        {
          "id": 1029866,
          "postDate": "2020-09-28T07:43:42.967Z",
          "content": "<p>Thanks, I guess we need to make some model design \\ post-processing to prevent this…</p>",
          "rawMarkdown": "Thanks, I guess we need to make some model design \\ post-processing to prevent this..."
        }
      ]
    },
    {
      "id": 1014326,
      "postDate": "2020-09-17T11:02:03.027Z",
      "content": "<p>I tried <a href=\"https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements\" target=\"_blank\">setting up a notebook</a> that checks for the label consistency requirements listed here to the degree with which I understand them. It would be great to get some feedback from the organizers to see if this covers the provided consistency rules.</p>",
      "rawMarkdown": "I tried [setting up a notebook](https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements) that checks for the label consistency requirements listed here to the degree with which I understand them. It would be great to get some feedback from the organizers to see if this covers the provided consistency rules.",
      "replies": [
        {
          "id": 1014733,
          "postDate": "2020-09-17T16:55:58.720Z",
          "content": "<p>Thank you for your work on this. I haven't tested the code, but your logical statement of the rules to be enforced appears to be correct and comprehensive. I will work on testing the code over the next several days and will post an update if I discover bugs.</p>",
          "rawMarkdown": "Thank you for your work on this. I haven't tested the code, but your logical statement of the rules to be enforced appears to be correct and comprehensive. I will work on testing the code over the next several days and will post an update if I discover bugs.",
          "votes": 1
        },
        {
          "id": 1014882,
          "postDate": "2020-09-17T18:55:31.647Z",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> great, thanks for looking into this!</p>",
          "rawMarkdown": "@anthracene great, thanks for looking into this!"
        },
        {
          "id": 1015126,
          "postDate": "2020-09-18T01:06:54.807Z",
          "content": "<p>[This comment deleted by author.]. </p>",
          "rawMarkdown": "[This comment deleted by author.]. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1013656,
      "postDate": "2020-09-16T21:11:13.330Z",
      "content": "<p>Thanks for clearing that up. Are those rules now already exhaustive? If not, what about cases, where the probabilities add up to more than 1? E.g. negative = 0.4, indeterminate = 0.4 and (Positive) = 0.4 (for at least one slice), or similar?</p>\n<p>If those examples are not exhaustive, would it be possible to provide a small script that checks whether a submission does fulfill the requirement (in the cases you listed this would be easy enough).</p>",
      "rawMarkdown": "Thanks for clearing that up. Are those rules now already exhaustive? If not, what about cases, where the probabilities add up to more than 1? E.g. negative = 0.4, indeterminate = 0.4 and (Positive) = 0.4 (for at least one slice), or similar?\n\nIf those examples are not exhaustive, would it be possible to provide a small script that checks whether a submission does fulfill the requirement (in the cases you listed this would be easy enough).",
      "replies": [
        {
          "id": 1013728,
          "postDate": "2020-09-16T22:55:20.410Z",
          "content": "<p>I believe that would be a case covered above: \"Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5\"</p>",
          "rawMarkdown": "I believe that would be a case covered above: \"Similarly, if no image is positive (p > 0.5), then there must be one and only one negative or indeterminate with p > 0.5\""
        },
        {
          "id": 1013986,
          "postDate": "2020-09-17T05:40:08.150Z",
          "content": "<p>Ok, then let's have (positive) &gt; 0.5, then we still have that the sum over negative, indeterminate, and (positive) is larger than 1 - would that be okay then? </p>",
          "rawMarkdown": "Ok, then let's have (positive) > 0.5, then we still have that the sum over negative, indeterminate, and (positive) is larger than 1 - would that be okay then? "
        },
        {
          "id": 1014754,
          "postDate": "2020-09-17T17:15:23.243Z",
          "content": "<p>We will not be enforcing any rules based on summations of probabilities, just on inconsistent combinations of probabilities greater than 0.5. </p>",
          "rawMarkdown": "We will not be enforcing any rules based on summations of probabilities, just on inconsistent combinations of probabilities greater than 0.5. ",
          "votes": 2
        },
        {
          "id": 1014806,
          "postDate": "2020-09-17T17:56:25.903Z",
          "content": "<p>Great, thanks for the clarification!</p>",
          "rawMarkdown": "Great, thanks for the clarification!"
        }
      ]
    },
    {
      "id": 1062423,
      "postDate": "2020-10-27T20:06:11.550Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1057723,
      "postDate": "2020-10-22T23:32:36.297Z",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks\n    "
    }
  ],
  "comments": [
    {
      "id": 1013821,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-09-17T02:23:07.643000",
      "content": "<p>To avoid a fiasco where one of the winners ends up getting disqualified because of a single prediction that ends up violating these requirements, it should be built into the submission system to throw an error if these requirements are not met. </p>\n<p>That would be the kind thing to do. </p>",
      "votes": 13,
      "replies": [
        {
          "id": 1013955,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-09-17T05:05:28.860000",
          "content": "<p>Agreed. <br>\nThere is always a chance of misunderstanding/not noticing these requirements. Would be very nice if it's implemented at kaggle side.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1014729,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-09-17T16:53:27.993000",
          "content": "<p>Automated checking of all submissions would be ideal, but unfortunately wasn't logistically feasible for this competition. We will be checking the winning submissions for consistency. The goal is not to trip people up; we don't want to disqualify anyone. At the same time, given the complexity of the metric, if logically inconsistent outputs are discovered that are scored by the metric, we don't want to be in the position of awarding algorithms that produce nonsensical output as winners. We will be working to make sure that the script that will be used to validate consistency of the winners is available to all contestants during the competition so they can check their own outputs for consistency (see <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> notebook later in this thread). We recognize this is less convenient than automated checking, but we hope this will reduce the chances of unhappy surprises.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1014525,
      "author_name": "quadcore/Richard Epstein",
      "author_url": "",
      "post_date": "2020-09-17T14:03:35.490000",
      "content": "<p>These consistency requirements, while problematic from the modeling viewpoint, are important from the perspective of \"Explainable AI\". </p>\n<p>Many \"models\" give impressive statistical results but are very difficult to move into actual clinical practice.</p>\n<p>For a model to be useful, it must be medically consistent, otherwise the physicians/health care providers will not trust it.</p>\n<p>For instance, if you report 8 images with 99% probably of PE, but your overall score for the exam is 10%, what is the clinician to think? What is their next action. Treat? Ignore?</p>\n<p>Do I treat a 30% probably of PE? At 50%? 80%?</p>\n<p>A prediction of 40% Positive, 30% Negative and 30% Indeterminate gives no guidance to treatment.</p>\n<p>We have all seen heat maps where the model predicts \"Pneumonia\" and the heat map is on the Right/Left marker.</p>\n<p>This competition is not asking for bounding boxes, but that is essentially what Right/Left/Central are. It doesn't help to say \"Positive for Pulmonary Embolism\" but not be able to show the clinician where it is.</p>\n<p>RV/LV ratio is essentially a whole separate Competition hidden within this one. RV/LV ratio can be elevated in the setting of pulmonary embolism and can signify the need for more intensive monitoring and treatment. It really is a binary classifier because by definition either the RV/LV is greater or less than one. The challenge for the model is first you have to find the heart and then you have to determine the RV/LV ratio. Most images might contribute nothing because they don't include the heart.</p>\n<p>While we have some work to do, \"Explainability\" and medical consistency are crucial features to get these model accepted and used.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1016032,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-09-18T15:59:46.777000",
      "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> you stated above <code>Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs.</code> But as the metric give 0 weight to the ‘PE Present on Image’ of images from series with no PE. Hence the \"punishment\" of the metric for predicting 'Negative for PE =1' and ‘PE Present on Image=1’ for images in studies with no PE is 0! </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1024457,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-09-23T22:05:30.893000",
      "content": "<p>Please take a look, as the host has added an officially-vetted version of <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">the code</a> that will be used to check compliance with these requirements, appended to the end of this original post.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1035686,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2020-10-03T00:33:52.007000",
          "content": "<p>In the code, the rule, <code>If any image is predicted positive, there cannot also be a predicted probability of Negative &gt; 0.5 nor can there be a predicted probability of Indeterminate &gt; 0.5.</code>,  is not checked., doesn't it?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 1035697,
              "author_name": "quadcore/Richard Epstein",
              "author_url": "",
              "post_date": "2020-10-03T01:03:47.167000",
              "content": "<p>Our understanding is that right now, none of the consistency rules are checked or enforced on the leaderboard. As discussed above, they plan on checking the top ten on the Private Leaderboard at the end of the competition. Others have posted Notebooks that implement the described checks.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 1035850,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-10-03T06:40:05.873000",
          "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> looks like you are correct this condition isn't check in the released code. <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> could you please take a look.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1035861,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-10-03T06:59:43.027000",
          "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> you are right, this rule was missing. I implemented it in <a href=\"https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements?scriptVersionId=43935750\" target=\"_blank\">Version 5</a> of my consistency check notebook. </p>\n<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> if it looks fine to you, maybe you could update the official version as well.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1041777,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-10-07T23:03:37.793000",
          "content": "<p>Thank you; we've updated the official version to add enforcement of this rule that was missing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1013864,
      "author_name": "yama",
      "author_url": "",
      "post_date": "2020-09-17T03:13:40.027000",
      "content": "<blockquote>\n  <p>Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5<br>\n  When no images are predicted positive, then none of these labels may be assigned p &gt; 0.5</p>\n</blockquote>\n<p>I am confused.  positive prob &lt; 0.5 for all images does not mean that exam level negative prob &gt; 0.5.</p>\n<p>Data desc page says that:</p>\n<blockquote>\n  <p>negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.</p>\n</blockquote>\n<p>So   <code>[Prob of Exam level Negative] = Π [prob that i-th image is nagative] = Π(1 - P_positive_i)</code>.<br>\nThen <code>P_positive_i &lt; 0.5 for all i-th image</code> does not mean <code>[Prob of Exam level Negative] &gt; 0.5</code><br>\nFor example, <code>P_positive = 0.4 for all i and number of image in exam is 2</code>, then <code>[Prob of Exam level Negative] = 0.6*0.6=0.36 &lt; 0.5</code></p>\n<p>Could you give me some explanation?</p>\n<p>[Update]<br>\nI also think below is not valid requirement.</p>\n<blockquote>\n  <p>Right, left, central -- if any image is predicted positive (p &gt; 0.5) then at least one of these labels must be assigned p &gt; 0.5</p>\n</blockquote>\n<p>Forecasting tomorrow's weather as \"30% shiny, 30% snowy, 40% rainy\" is definitely valid prediction. If you want at least one positive out of 3 candidates, you can use threshold of 1/3 instead of 1/2. There exists no reason to use 1/2.  <br>\nIt is ok to require <code>P(left)+P(central)+P(right) &gt;= 1</code> and use threthold of 1/3 to get at least one postive.  But host requires more</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1014743,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-09-17T17:01:18.960000",
          "content": "<p><a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> 's post in this thread does an excellent job of explaining the rationale for these rules</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1015480,
          "author_name": "yama",
          "author_url": "",
          "post_date": "2020-09-18T08:06:33.863000",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> Thanks for your resnponce.</p>\n<p>What host wants can be summarized as \"at least one positive\".<br>\nThat can be easily implemented in the application-phase. The apps does not even have to show probabilities to users.<br>\nI believe that modelling-phase should not have a responsibility for it.</p>\n<p>Especially, logloss metrics implicitly assumes calibrated probability as a prediction. Statistically meaningless requirements distort the evaluation.<br>\nI think that the requirement should be amended to make this competition better.</p>\n<p>I admit (some) users may feel nervous to \"30% left, 30% central, 40% right\" prediction.<br>\nBut \"at least 50% prob requirements\" does not resolve the issue and it is not related to \"logical consistency\". There is no difference between 30%, 33.3% and 50%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1050925,
          "author_name": "DarkCube",
          "author_url": "",
          "post_date": "2020-10-15T23:03:14.890000",
          "content": "<p>[Prob of Exam level Negative] =/= Π [prob that i-th image is nagative] !!!!<br>\nΠ [prob that i-th image is nagative] = is the probability that <em>all</em> images are negative for PE.<br>\na better approximation for [Prob of Exam level Negative] is max probability that  i-th image is nagative.</p>\n<p>Moreover, there can be more than one PE in a single lung. Therefore, Right, left, central are NOT mutually exclusive (e.g. there can be a right AND a central PE in the same lung).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1058230,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-23T13:31:24.130000",
      "content": "<p>Good, but what about the training set ? There are such inconsistencies in it too, for example cases when neither of <code>pe_present_on_image, negative_exam_for_pe, indeterminate</code> are equal to one. Perhaps this has already been mentioned in discussions or notebooks, and I missed it. If it is the case, I apologize.</p>\n<p>I used this code : </p>\n<pre><code>labels = ['pe_present_on_image', 'negative_exam_for_pe', 'indeterminate']\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n</code></pre>\n<p>We obtain : </p>\n<p>Total nb of images : 1790594<br>\nImages with ONE of the three categories mentioned above, equal to one : 1344365<br>\nImages with NONE of the three categories mentioned above, equal to one : 446229</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1058277,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-23T14:25:49.810000",
          "content": "<p>You need to do this. For each study/exam, you need to calculate <code>max_pe_present</code> in order to know whether that study is positive or not.</p>\n<pre><code>labels = ['max_pe_present', 'negative_exam_for_pe', 'indeterminate']\ndf_train['max_pe_present'] = \\   \n  df_train.groupby('StudyInstanceUID').pe_present_on_image.transform('max')\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n</code></pre>\n<p>Then it prints 1790594 and 1790594</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1058285,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-23T14:33:11.613000",
          "content": "<p>Indeed, you are right. Thank you. Of course, we need to have at least one image per study with   pe_present_on_image=1 in order to have a positive study.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1054463,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-20T00:40:43.077000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> </p>\n<p>The rules in this post (and the competition rules <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/data\" target=\"_blank\">here</a>) do not agree with the rules in the official notebook <a href=\"https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\" target=\"_blank\">here</a>. Could you please clarify which is correct?</p>\n<p>Specifically, this thread allows us to predict both <code>pe_present_on_image&lt;=0.5</code> and <code>rv_lv_ratio_gte_1 &gt; 0.5</code> for an entire exam/study whereas the notebook does not allow this.</p>\n<p>Can a patient have a right ventricle that is larger than a left ventricle and still not have pulmonary embolism?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1055349,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-10-20T17:26:28.260000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for raising! The host is reviewing your feedback and will respond directly as soon as possible.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1057644,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-22T20:21:10.243000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , the host responded that if exam has all <code>pe_present_on_image&lt;=0.5</code> then <code>rv_lv_ratio_gte_1</code> and <code>rv_lv_ratio_lt_1</code> must be <code>&lt;=0.5</code>.</p>\n<p>I have another question. Which teams' submissions will be checked for Label Consistency? And what if one submission conforms but the second submission does not conform?</p>\n<p>I notice the <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/rules\" target=\"_blank\">rules</a> say below but I was wondering if you could be more specific like (1) only cash prizes will be checked (2) or all medal positions will be checked, etc.</p>\n<blockquote>\n  <p>Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page. This includes:</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057647,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-22T20:22:35.970000",
          "content": "<p>Thanks for your response in my other <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696\" target=\"_blank\">deleted post</a>. You said</p>\n<blockquote>\n  <p><strong>Prize winners'</strong> winning submissions will be checked and disqualified/removed if non-conforming. They should not expect to rely on any secondary selected submission as a fall-back.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057725,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-10-22T23:35:52.860000",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <br>\nI'd like to know too.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057728,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-22T23:43:41.697000",
          "content": "<p>Both questions have been answered. The host answered about <code>rv_lv_ratio_gte_1</code> below. And Julia answered about leaderboard removal <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/192696\" target=\"_blank\">here</a>. (Originally i asked the question in a separate topic but moved it here. I deleted it before i noticed that Julia responded).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057755,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-10-23T00:56:22.210000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057756,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-10-23T00:57:36.197000",
          "content": "<p>oh sry I missed it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1052550,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-17T22:16:19.553000",
      "content": "<p>From the explanation and from the notebook which implements these errors' checking, it transpires that for an exam with at least one positive image, it is allowed to have both Acute &lt;= 0.5 and Acute and Chronic &lt;= 0.5. </p>\n<p>On the other hand, in the description of the competition (the schema) it is written that for positive exam we need to have at least one of Acute or Acute and Chronic &gt; 0.5.</p>\n<p>Which statement is true ? Can someone explain ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052564,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-17T23:31:17.297000",
          "content": "<p>The schema below only uses the words \"at least one\" for \"right-sided\", \"left-sided\", and \"central\". Regarding \"RV/LV\" targets and \"Chronic Acute\" targets. It says \"only one\" which means \"at most one which can be zero\".</p>\n<p>To answer your answer, if you have a positive exam, you are allowed to leave both \"chronic\" and \"chronic+acute\" &lt;=0.5 for that exam. (This represents the case of just \"acute\" for that exam)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F69f7671bbaa31bbda88224e4d16a509f%2Fscheme.png?generation=1602977187941229&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1052567,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-17T23:35:43.283000",
          "content": "<p>Thank you! The truth is it is a little ambiguous in the schema. It is clear now. And simplifies things.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1052538,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-17T21:28:45.770000",
      "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> <a href=\"https://www.kaggle.com/jeffrudie\" target=\"_blank\">@jeffrudie</a> </p>\n<p>John, does your linked notebook (latest version 2) enforce the rules (what you've written above) correctly? The notebook says that we cannot predict <code>rv_lv_ratio_lt_1 &gt;0.5</code> nor <code>rv_lv_ratio_gte_1 &gt;0.5</code> if all <code>pe_present_on_image&lt;=0.5</code> for a given study with the following code</p>\n<pre><code>rule2b = df_neg.loc[(df_neg.rv_lv_ratio_lt_1     &gt; 0.5) | \n                (df_neg.rv_lv_ratio_gte_1    &gt; 0.5) |\n                (df_neg.central_pe           &gt; 0.5) | \n                (df_neg.rightsided_pe        &gt; 0.5) | \n                (df_neg.leftsided_pe         &gt; 0.5) |\n                (df_neg.acute_and_chronic_pe &gt; 0.5) | \n                (df_neg.chronic_pe           &gt; 0.5)].reset_index(drop = True)\n</code></pre>\n<p>I don't think this agrees with what you wrote above</p>\n<blockquote>\n  <p>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive.<br>\n  if any image on the exam is positive, one of these must have p &gt; 0.5<br>\n  both cannot have p &gt; 0.5</p>\n</blockquote>\n<p>The statement you wrote allows us to predict <code>rv_lv_ratio_lt_1 &gt;0.5</code> or <code>rv_lv_ratio_gte_1 &gt;0.5</code> if all <code>pe_present_on_image&lt;=0.5</code> for a given study. </p>\n<p>Am i correct that your code is different that your text, and if so, which rule is correct?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052540,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-17T21:34:33.263000",
          "content": "<p>IMHO, you need to change your notebook to the following if you want to agree with what you wrote above</p>\n<pre><code>    rule2b = df_neg.loc[((df_neg.rv_lv_ratio_lt_1     &gt; 0.5) &amp; \n                    (df_neg.rv_lv_ratio_gte_1    &gt; 0.5)) |\n                    (df_neg.central_pe           &gt; 0.5) | \n                    (df_neg.rightsided_pe        &gt; 0.5) | \n                    (df_neg.leftsided_pe         &gt; 0.5) |\n                    (df_neg.acute_and_chronic_pe &gt; 0.5) | \n                    (df_neg.chronic_pe           &gt; 0.5)].reset_index(drop = True)\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057551,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-10-22T18:43:25.160000",
          "content": "<p>I believe the code in the notebook is correct; my English description of the rule may not be completely clear. If the study is negative, then neither rv_lv_ratio_lt_1 nor rv_lv_ratio_gte_1 may be predicted as &gt;0.5.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057576,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-10-22T19:11:10.620000",
          "content": "<p>Ok, thanks for the update</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1049766,
      "author_name": "James Ingram",
      "author_url": "",
      "post_date": "2020-10-14T18:58:58.657000",
      "content": "<p>Excellent work!</p>\n<p>btw:</p>\n<ul>\n<li>v3: wrapped consistency checks into function</li>\n</ul>\n<p>should be:<br>\ncheck_consistency()</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047698,
      "author_name": "Catadanna",
      "author_url": "",
      "post_date": "2020-10-12T21:22:02.147000",
      "content": "<p>Proud to get this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3616012%2F2689841674e17fa9bf6a21c1f041a783%2FRSNA_no_errors.png?generation=1602537705968718&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1047977,
          "author_name": "SumanSudhir",
          "author_url": "",
          "post_date": "2020-10-13T05:20:45.760000",
          "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> what about private test data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048091,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-13T07:20:28.060000",
          "content": "<p>What do you mean by that? This is the result obtained on one of my submissions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048181,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2020-10-13T09:00:36.040000",
          "content": "<p>The reply is on the public test. You can't see the reply for the private test.</p>\n<p>What I do for the private test is something like:</p>\n<pre><code>If len(errors)==0:\n    save submission file\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1048196,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2020-10-13T09:17:56.190000",
          "content": "<p><a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a>: Indeed, that is what I thought too! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1024412,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-09-23T20:53:02.867000",
      "content": "<p>Are these rules currently enforced in the submission grader? I had submission scoring errors submitting the sample_submission (all probabilities at 0.5) and I wonder if this is the reason?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1024452,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-09-23T21:58:40.050000",
          "content": "<p>No, unfortunately they aren't. We will manually run the top 10 private leaderboard entries against this code at the conclusion of the contest to verify compliance.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1029788,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-28T06:13:52.423000",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> Sorry if this has been asked before or due to my misunderstanding of the competition rule. May I ask what will happen to the top-10 if they fail the compliance check? (No offense to the current top-10!)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1029834,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-09-28T07:14:09.147000",
          "content": "<p>Quoting the competition rules:</p>\n<blockquote>\n  <p>Competition Sponsor reserves the right to disqualify any participant who makes conflicting label predictions that do not adhere to the expected label hierarchy defined by the diagram on the Data page.</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>A disqualified participant will be removed from the Competition leaderboard, at the Competition Sponsor and Kaggle's sole discretion. If a Participant is removed from the Competition Leaderboard, additional winning features associated with the Kaggle competition platform, for example Kaggle points or medals, may also not be awarded.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1029866,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2020-09-28T07:43:42.967000",
          "content": "<p>Thanks, I guess we need to make some model design \\ post-processing to prevent this…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1014326,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-09-17T11:02:03.027000",
      "content": "<p>I tried <a href=\"https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements\" target=\"_blank\">setting up a notebook</a> that checks for the label consistency requirements listed here to the degree with which I understand them. It would be great to get some feedback from the organizers to see if this covers the provided consistency rules.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1014733,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-09-17T16:55:58.720000",
          "content": "<p>Thank you for your work on this. I haven't tested the code, but your logical statement of the rules to be enforced appears to be correct and comprehensive. I will work on testing the code over the next several days and will post an update if I discover bugs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1014882,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2020-09-17T18:55:31.647000",
          "content": "<p><a href=\"https://www.kaggle.com/anthracene\" target=\"_blank\">@anthracene</a> great, thanks for looking into this!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1015126,
          "author_name": "ThePig",
          "author_url": "",
          "post_date": "2020-09-18T01:06:54.807000",
          "content": "<p>[This comment deleted by author.]. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1013656,
      "author_name": "maettes",
      "author_url": "",
      "post_date": "2020-09-16T21:11:13.330000",
      "content": "<p>Thanks for clearing that up. Are those rules now already exhaustive? If not, what about cases, where the probabilities add up to more than 1? E.g. negative = 0.4, indeterminate = 0.4 and (Positive) = 0.4 (for at least one slice), or similar?</p>\n<p>If those examples are not exhaustive, would it be possible to provide a small script that checks whether a submission does fulfill the requirement (in the cases you listed this would be easy enough).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1013728,
          "author_name": "Robyn Ball",
          "author_url": "",
          "post_date": "2020-09-16T22:55:20.410000",
          "content": "<p>I believe that would be a case covered above: \"Similarly, if no image is positive (p &gt; 0.5), then there must be one and only one negative or indeterminate with p &gt; 0.5\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1013986,
          "author_name": "maettes",
          "author_url": "",
          "post_date": "2020-09-17T05:40:08.150000",
          "content": "<p>Ok, then let's have (positive) &gt; 0.5, then we still have that the sum over negative, indeterminate, and (positive) is larger than 1 - would that be okay then? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1014754,
          "author_name": "John Mongan",
          "author_url": "",
          "post_date": "2020-09-17T17:15:23.243000",
          "content": "<p>We will not be enforcing any rules based on summations of probabilities, just on inconsistent combinations of probabilities greater than 0.5. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1014806,
          "author_name": "maettes",
          "author_url": "",
          "post_date": "2020-09-17T17:56:25.903000",
          "content": "<p>Great, thanks for the clarification!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1062423,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T20:06:11.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1057723,
      "author_name": "YujiAriyasu",
      "author_url": "",
      "post_date": "2020-10-22T23:32:36.297000",
      "content": "<p>thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1013615": "The data page for this competition describes a requirement for logical consistency of predicted labels, and includes a diagram describing the relationship between labels. \n\nFor the purpose of enforcing this rule we consider a label to be “predicted” if that label is assigned a probability of > 0.5. We require that the “predicted” labels be logically consistent. Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs. However, due to the complexity of the metric and the difficulty in anticipating all corner cases, we have put this rule in place to reduce the ability to game the metric with nonsensical outputs.\n\nSome specific examples of how this rule plays out:\n\nAt the image level, any image with predicted probability >  0.5 is considered as being positive for PE will count as a positive image\n\nAt the exam level, we have\n1. Negative, Indeterminate, (Positive) and it can only be one of these. If any image is predicted positive, there cannot also be a predicted probability of Negative > 0.5 nor can there be a predicted probability of Indeterminate >  0.5.\n\nSimilarly, if no image is positive (p > 0.5), then there must be one and only one negative or indeterminate with p > 0.5\n\n2. Right, left, central -- if any image is predicted positive (p > 0.5) then at least one of these labels must be assigned p > 0.5; more than one of these labels may be assigned p > 0.5. When no images are predicted positive, then none of these labels may be assigned p > 0.5\n\n3. RV/LV ratio. It can be only one of these and it must be present if at least one image is positive. \n- if any image on the exam is positive, one of these must have p >  0.5\n- both cannot have p >  0.5\n\n4. Acute, Chronic, Acute & Chronic -- it cannot be both chronic & acute and chronic so \n- only one can have p >  0.5\n- it is also possible that neither has p > 0.5  \n- in other words, it is inconsistent to say chronic has p > 0.5 and acute & chronic has p >  0.5.\n\n**EDIT:** The code that will be used to check compliance with these requirements is available in this [notebook](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check).",
    "1013821": "To avoid a fiasco where one of the winners ends up getting disqualified because of a single prediction that ends up violating these requirements, it should be built into the submission system to throw an error if these requirements are not met. \n\nThat would be the kind thing to do. \n\n",
    "1014525": "These consistency requirements, while problematic from the modeling viewpoint, are important from the perspective of \"Explainable AI\". \n\nMany \"models\" give impressive statistical results but are very difficult to move into actual clinical practice.\n\nFor a model to be useful, it must be medically consistent, otherwise the physicians/health care providers will not trust it.\n\nFor instance, if you report 8 images with 99% probably of PE, but your overall score for the exam is 10%, what is the clinician to think? What is their next action. Treat? Ignore?\n\nDo I treat a 30% probably of PE? At 50%? 80%?\n\nA prediction of 40% Positive, 30% Negative and 30% Indeterminate gives no guidance to treatment.\n\nWe have all seen heat maps where the model predicts \"Pneumonia\" and the heat map is on the Right/Left marker.\n\nThis competition is not asking for bounding boxes, but that is essentially what Right/Left/Central are. It doesn't help to say \"Positive for Pulmonary Embolism\" but not be able to show the clinician where it is.\n\nRV/LV ratio is essentially a whole separate Competition hidden within this one. RV/LV ratio can be elevated in the setting of pulmonary embolism and can signify the need for more intensive monitoring and treatment. It really is a binary classifier because by definition either the RV/LV is greater or less than one. The challenge for the model is first you have to find the heart and then you have to determine the RV/LV ratio. Most images might contribute nothing because they don't include the heart.\n\nWhile we have some work to do, \"Explainability\" and medical consistency are crucial features to get these model accepted and used.",
    "1016032": "@anthracene you stated above `Note that we expect that logically inconsistent outputs will score poorly on the metric, and that high-scoring algorithms will naturally produce logically consistent outputs.` But as the metric give 0 weight to the ‘PE Present on Image’ of images from series with no PE. Hence the \"punishment\" of the metric for predicting 'Negative for PE =1' and ‘PE Present on Image=1’ for images in studies with no PE is 0! ",
    "1024457": "Please take a look, as the host has added an officially-vetted version of [the code](https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check) that will be used to check compliance with these requirements, appended to the end of this original post.",
    "1013864": "> Similarly, if no image is positive (p > 0.5), then there must be one and only one negative or indeterminate with p > 0.5\n> When no images are predicted positive, then none of these labels may be assigned p > 0.5\n\nI am confused.  positive prob < 0.5 for all images does not mean that exam level negative prob > 0.5.\n\nData desc page says that:\n> negative_exam_for_pe - exam-level, whether there are any images in the study that have PE present.\n\nSo   `[Prob of Exam level Negative] = Π [prob that i-th image is nagative] = Π(1 - P_positive_i)`.\nThen `P_positive_i < 0.5 for all i-th image` does not mean `[Prob of Exam level Negative] > 0.5`\nFor example, `P_positive = 0.4 for all i and number of image in exam is 2`, then `[Prob of Exam level Negative] = 0.6*0.6=0.36 < 0.5`\n\nCould you give me some explanation?\n\n\n[Update]\nI also think below is not valid requirement.\n> Right, left, central -- if any image is predicted positive (p > 0.5) then at least one of these labels must be assigned p > 0.5\n\nForecasting tomorrow's weather as \"30% shiny, 30% snowy, 40% rainy\" is definitely valid prediction. If you want at least one positive out of 3 candidates, you can use threshold of 1/3 instead of 1/2. There exists no reason to use 1/2.  \nIt is ok to require `P(left)+P(central)+P(right) >= 1` and use threthold of 1/3 to get at least one postive.  But host requires more",
    "1058230": "Good, but what about the training set ? There are such inconsistencies in it too, for example cases when neither of `pe_present_on_image, negative_exam_for_pe, indeterminate` are equal to one. Perhaps this has already been mentioned in discussions or notebooks, and I missed it. If it is the case, I apologize.\n\nI used this code : \n\n```\nlabels = ['pe_present_on_image', 'negative_exam_for_pe', 'indeterminate']\ndf_train['number_level_one_categories'] = df_train[labels].sum(axis=1)\nprint(df_train.shape[0])\ndf_train_1 = df_train.loc[df_train['number_level_one_categories']==1]\nprint(df_train_1.shape[0])\n```\n\nWe obtain : \n\nTotal nb of images : 1790594\nImages with ONE of the three categories mentioned above, equal to one : 1344365\nImages with NONE of the three categories mentioned above, equal to one : 446229",
    "1054463": "Hello @anthracene @philculliton @juliaelliott \n\nThe rules in this post (and the competition rules [here][2]) do not agree with the rules in the official notebook [here][1]. Could you please clarify which is correct?\n\nSpecifically, this thread allows us to predict both `pe_present_on_image<=0.5` and `rv_lv_ratio_gte_1 > 0.5` for an entire exam/study whereas the notebook does not allow this.\n\nCan a patient have a right ventricle that is larger than a left ventricle and still not have pulmonary embolism?\n\n[1]: https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check\n[2]: https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/data",
    "1052550": "From the explanation and from the notebook which implements these errors' checking, it transpires that for an exam with at least one positive image, it is allowed to have both Acute <= 0.5 and Acute and Chronic <= 0.5. \n\nOn the other hand, in the description of the competition (the schema) it is written that for positive exam we need to have at least one of Acute or Acute and Chronic > 0.5.\n\nWhich statement is true ? Can someone explain ? ",
    "1052538": "@anthracene @jeffrudie \n\nJohn, does your linked notebook (latest version 2) enforce the rules (what you've written above) correctly? The notebook says that we cannot predict `rv_lv_ratio_lt_1 >0.5` nor `rv_lv_ratio_gte_1 >0.5` if all `pe_present_on_image<=0.5` for a given study with the following code\n  \n    rule2b = df_neg.loc[(df_neg.rv_lv_ratio_lt_1     > 0.5) | \n                    (df_neg.rv_lv_ratio_gte_1    > 0.5) |\n                    (df_neg.central_pe           > 0.5) | \n                    (df_neg.rightsided_pe        > 0.5) | \n                    (df_neg.leftsided_pe         > 0.5) |\n                    (df_neg.acute_and_chronic_pe > 0.5) | \n                    (df_neg.chronic_pe           > 0.5)].reset_index(drop = True)\n\nI don't think this agrees with what you wrote above\n\n>RV/LV ratio. It can be only one of these and it must be present if at least one image is positive.\nif any image on the exam is positive, one of these must have p > 0.5\nboth cannot have p > 0.5\n\nThe statement you wrote allows us to predict `rv_lv_ratio_lt_1 >0.5` or `rv_lv_ratio_gte_1 >0.5` if all `pe_present_on_image<=0.5` for a given study. \n\nAm i correct that your code is different that your text, and if so, which rule is correct?\n\n[1]: https://www.kaggle.com/anthracene/host-confirmed-label-consistency-check",
    "1049766": "Excellent work!\n\nbtw:\n\n* v3: wrapped consistency checks into~~ check_consitency()~~ function\n\nshould be:\ncheck_consistency()",
    "1047698": "Proud to get this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3616012%2F2689841674e17fa9bf6a21c1f041a783%2FRSNA_no_errors.png?generation=1602537705968718&alt=media)",
    "1024412": "Are these rules currently enforced in the submission grader? I had submission scoring errors submitting the sample_submission (all probabilities at 0.5) and I wonder if this is the reason?",
    "1014326": "I tried [setting up a notebook](https://www.kaggle.com/kozodoi/checking-the-label-consistency-requirements) that checks for the label consistency requirements listed here to the degree with which I understand them. It would be great to get some feedback from the organizers to see if this covers the provided consistency rules.",
    "1013656": "Thanks for clearing that up. Are those rules now already exhaustive? If not, what about cases, where the probabilities add up to more than 1? E.g. negative = 0.4, indeterminate = 0.4 and (Positive) = 0.4 (for at least one slice), or similar?\n\nIf those examples are not exhaustive, would it be possible to provide a small script that checks whether a submission does fulfill the requirement (in the cases you listed this would be easy enough).",
    "1062423": "",
    "1057723": "thanks\n    "
  }
}