{
  "id": 193417,
  "title": "9th place solution ( + github code)",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193417",
  "author_name": "shimacos",
  "post_date": "2020-10-27T01:25:18.906000",
  "votes": 21,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations to all winners !<br>\nThis competition was hard on me in many ways.</p>\n<h1>Solution Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fb7889dcd4b8229c53e2103cfb622a8e1%2FRSNA%202020%20solution%20(1).png?generation=1603760751579416&amp;alt=media\" alt=\"\"></p>\n<h1>Preprocess</h1>\n<ul>\n<li><p>In the train data, no CT image have PE after 400th image. So, we used only images before 400th image.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fc51f745506070907901a148f396ae5d7%2Fdownload-20.png?generation=1603800720901371&amp;alt=media\" alt=\"\"></p></li>\n<li><p>For stage 1 training, we preprocessed image-level labels like following image.</p>\n<ul>\n<li>Before<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F41ffbb129012a32e592556611c4a155d%2Fdownload-21.png?generation=1603760930759265&amp;alt=media\" alt=\"\"></li>\n<li>After<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F0a0fa0aacd34ca23e8b7a00e385d25b0%2Fdownload-22.png?generation=1603760956442464&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h1>Stage 1 training</h1>\n<ul>\n<li>We used 512 x 512 image + efficientnet-b5 and 384 x 384 image + efficientnet-b3 and using preprocessed labels.</li>\n</ul>\n<h1>Stage 2 training</h1>\n<ul>\n<li>Inference time was so severe because we used 512 x 512 image + efficientnet-b5. So, we subsampled 400 sequences to 200 sequences and used Deconvolution module.<ul>\n<li>We got the same CV score when using 400 sequences.</li></ul></li>\n<li>We was not able to use various models in stage 1 because of resource. Therefore, we trained various models in stage 2.<ul>\n<li>Input: b5-feature only, b3-feature only, b5-feature + b3-feature</li>\n<li>model: Conv1D, LSTM, GRU, Conv1D + LSTM</li>\n<li>output: 3 x 4 = 12 predictions</li></ul></li>\n</ul>\n<h1>Stacking</h1>\n<ul>\n<li>We trained LGBM, Conv1D and GRU.<ul>\n<li>We used only PE-exam when training pe_present_on_image by lgbm because image from negative PE doesn't affect competition metric.</li></ul></li>\n</ul>\n<h1>Postprocess</h1>\n<ul>\n<li>We've implemented a heuristic post process.<ul>\n<li>This postprocess increased CV and public score 0.002 (Private 0.160 -&gt; 0.162)</li>\n<li>Main idea<ul>\n<li>replace  <code>pe_present_on_image</code> with <code>1 - negative_exam_for_pe</code> when <code>1 - negative_exam_for_pe &lt;= pe_present_on_image</code></li>\n<li>repeat sigmoid -&gt; logit -&gt; logit += s -&gt; sigmoid until satisfying label consistency</li></ul></li></ul></li>\n</ul>\n<pre><code>label_cols = [\n      \"pe_present_on_image\",\n      \"negative_exam_for_pe\",\n      \"indeterminate\",\n      \"chronic_pe\",\n      \"acute_and_chronic_pe\",\n      \"central_pe\",\n      \"leftsided_pe\",\n      \"rightsided_pe\",\n      \"rv_lv_ratio_gte_1\",\n      \"rv_lv_ratio_lt_1\",\n    ]\n\ndef postprocess(x, s=2.0):\n    logit = np.log(x/(1 - x))\n    logit = logit + s\n    sigmoid = 1 / (1 + np.exp(-logit))\n    return sigmoid\n\ndef satisfy_label_consistency(df):\n    rule_breaks = consistency_check(df).index\n    print(rule_breaks)\n    if len(rule_breaks) &gt; 0:\n        df[\"positive_exam_for_pe\"] = 1 - df[\"negative_exam_for_pe\"]\n        df.loc[\n            df.query(\"positive_exam_for_pe &lt;= pe_present_on_image\").index,\n            \"pe_present_on_image\",\n        ] = df.loc[\n            df.query(\"positive_exam_for_pe &lt;= pe_present_on_image\").index,\n            \"positive_exam_for_pe\",\n        ]\n        rule_breaks = consistency_check(df).index\n        df[\"positive_images_in_exam\"] = df[\"StudyInstanceUID\"].map(\n            df.groupby([\"StudyInstanceUID\"])[\"pe_present_on_image\"].max()\n        )\n        df_pos = df.query(\"positive_images_in_exam &gt; 0.5\")\n        df_neg = df.query(\"positive_images_in_exam &lt;= 0.5\")\n        if \"1a\" in rule_breaks:\n            rv_filter = \"rv_lv_ratio_gte_1 &gt; 0.5 &amp; rv_lv_ratio_lt_1 &gt; 0.5\"\n            while len(df_pos.query(rv_filter)) &gt; 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_min\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].min(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_min\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_min\")[\n                            rv_col\n                        ].values,\n                        s=-0.1,\n                    )\n            rv_filter = \"rv_lv_ratio_gte_1 &lt;= 0.5 &amp; rv_lv_ratio_lt_1 &lt;= 0.5\"\n            while len(df_pos.query(rv_filter)) &gt; 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_max\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].max(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_max\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_max\")[\n                            rv_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_pos.index, label_cols[8:]] = df_pos[label_cols[8:]]\n        if \"1b\" in rule_breaks:\n            pe_filter = \" &amp; \".join([f\"{col} &lt;= 0.5\" for col in label_cols[5:8]])\n            while \"1b\" in consistency_check(df).index:\n                for col in label_cols[5:8]:\n                    df_pos.loc[df_pos.query(pe_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(pe_filter).index, col], s=0.1\n                    )\n                df.loc[df_pos.index, label_cols[5:8]] = df_pos[label_cols[5:8]].values\n        if \"1c\" in rule_breaks:\n            chronic_filter = \"chronic_pe &gt; 0.5 &amp; acute_and_chronic_pe &gt; 0.5\"\n            df_pos.loc[df_pos.query(chronic_filter).index, label_cols[3:5]] = softmax(\n                df_pos.query(chronic_filter)[label_cols[3:5]].values, axis=1\n            )\n            df.loc[df_pos.index, label_cols[3:5]] = df_pos[label_cols[3:5]]\n        if \"1d\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe &gt; 0.5 | indeterminate &gt; 0.5\"\n            while \"1d\" in consistency_check(df).index:\n                for col in label_cols[1:3]:\n                    df_pos.loc[df_pos.query(neg_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(neg_filter).index, col], s=-0.1\n                    )\n                df.loc[df_pos.index, label_cols[1:3]] = df_pos[label_cols[1:3]].values\n        if \"2a\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe &gt; 0.5 &amp; indeterminate &gt; 0.5\"\n            while len(df_neg.query(neg_filter)) &gt; 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_min\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].min(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_min\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_min\")[\n                            neg_col\n                        ].values,\n                        s=-0.1,\n                    )\n            neg_filter = \"negative_exam_for_pe &lt;= 0.5 &amp; indeterminate &lt;= 0.5\"\n            while len(df_neg.query(neg_filter)) &gt; 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_max\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].max(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_max\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_max\")[\n                            neg_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_neg.index, label_cols[1:3]] = df_neg[label_cols[1:3]]\n        if \"2b\" in rule_breaks:\n            while \"2b\" in consistency_check(df).index:\n                for col in label_cols[3:]:\n                    df_neg.loc[df_neg.query(f\"{col} &gt; 0.5\").index, col] = postprocess(\n                        df_neg.loc[df_neg.query(f\"{col} &gt; 0.5\").index, col], s=-0.1\n                    )\n                df.loc[df_neg.index, label_cols[3:]] = df_neg[label_cols[3:]].values\n    return df\n</code></pre>\n<p>Updated: <br>\nWe uploaded code on github (<a href=\"https://github.com/shimacos37/kaggle-rsna-2020-9th-solution)\" target=\"_blank\">https://github.com/shimacos37/kaggle-rsna-2020-9th-solution)</a>.</p>",
  "messages": [
    {
      "id": 1061392,
      "postDate": "2020-10-27T01:25:18.907Z",
      "content": "<p>Congratulations to all winners !<br>\nThis competition was hard on me in many ways.</p>\n<h1>Solution Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fb7889dcd4b8229c53e2103cfb622a8e1%2FRSNA%202020%20solution%20(1).png?generation=1603760751579416&amp;alt=media\" alt=\"\"></p>\n<h1>Preprocess</h1>\n<ul>\n<li><p>In the train data, no CT image have PE after 400th image. So, we used only images before 400th image.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fc51f745506070907901a148f396ae5d7%2Fdownload-20.png?generation=1603800720901371&amp;alt=media\" alt=\"\"></p></li>\n<li><p>For stage 1 training, we preprocessed image-level labels like following image.</p>\n<ul>\n<li>Before<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F41ffbb129012a32e592556611c4a155d%2Fdownload-21.png?generation=1603760930759265&amp;alt=media\" alt=\"\"></li>\n<li>After<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F0a0fa0aacd34ca23e8b7a00e385d25b0%2Fdownload-22.png?generation=1603760956442464&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h1>Stage 1 training</h1>\n<ul>\n<li>We used 512 x 512 image + efficientnet-b5 and 384 x 384 image + efficientnet-b3 and using preprocessed labels.</li>\n</ul>\n<h1>Stage 2 training</h1>\n<ul>\n<li>Inference time was so severe because we used 512 x 512 image + efficientnet-b5. So, we subsampled 400 sequences to 200 sequences and used Deconvolution module.<ul>\n<li>We got the same CV score when using 400 sequences.</li></ul></li>\n<li>We was not able to use various models in stage 1 because of resource. Therefore, we trained various models in stage 2.<ul>\n<li>Input: b5-feature only, b3-feature only, b5-feature + b3-feature</li>\n<li>model: Conv1D, LSTM, GRU, Conv1D + LSTM</li>\n<li>output: 3 x 4 = 12 predictions</li></ul></li>\n</ul>\n<h1>Stacking</h1>\n<ul>\n<li>We trained LGBM, Conv1D and GRU.<ul>\n<li>We used only PE-exam when training pe_present_on_image by lgbm because image from negative PE doesn't affect competition metric.</li></ul></li>\n</ul>\n<h1>Postprocess</h1>\n<ul>\n<li>We've implemented a heuristic post process.<ul>\n<li>This postprocess increased CV and public score 0.002 (Private 0.160 -&gt; 0.162)</li>\n<li>Main idea<ul>\n<li>replace  <code>pe_present_on_image</code> with <code>1 - negative_exam_for_pe</code> when <code>1 - negative_exam_for_pe &lt;= pe_present_on_image</code></li>\n<li>repeat sigmoid -&gt; logit -&gt; logit += s -&gt; sigmoid until satisfying label consistency</li></ul></li></ul></li>\n</ul>\n<pre><code>label_cols = [\n      \"pe_present_on_image\",\n      \"negative_exam_for_pe\",\n      \"indeterminate\",\n      \"chronic_pe\",\n      \"acute_and_chronic_pe\",\n      \"central_pe\",\n      \"leftsided_pe\",\n      \"rightsided_pe\",\n      \"rv_lv_ratio_gte_1\",\n      \"rv_lv_ratio_lt_1\",\n    ]\n\ndef postprocess(x, s=2.0):\n    logit = np.log(x/(1 - x))\n    logit = logit + s\n    sigmoid = 1 / (1 + np.exp(-logit))\n    return sigmoid\n\ndef satisfy_label_consistency(df):\n    rule_breaks = consistency_check(df).index\n    print(rule_breaks)\n    if len(rule_breaks) &gt; 0:\n        df[\"positive_exam_for_pe\"] = 1 - df[\"negative_exam_for_pe\"]\n        df.loc[\n            df.query(\"positive_exam_for_pe &lt;= pe_present_on_image\").index,\n            \"pe_present_on_image\",\n        ] = df.loc[\n            df.query(\"positive_exam_for_pe &lt;= pe_present_on_image\").index,\n            \"positive_exam_for_pe\",\n        ]\n        rule_breaks = consistency_check(df).index\n        df[\"positive_images_in_exam\"] = df[\"StudyInstanceUID\"].map(\n            df.groupby([\"StudyInstanceUID\"])[\"pe_present_on_image\"].max()\n        )\n        df_pos = df.query(\"positive_images_in_exam &gt; 0.5\")\n        df_neg = df.query(\"positive_images_in_exam &lt;= 0.5\")\n        if \"1a\" in rule_breaks:\n            rv_filter = \"rv_lv_ratio_gte_1 &gt; 0.5 &amp; rv_lv_ratio_lt_1 &gt; 0.5\"\n            while len(df_pos.query(rv_filter)) &gt; 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_min\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].min(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_min\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_min\")[\n                            rv_col\n                        ].values,\n                        s=-0.1,\n                    )\n            rv_filter = \"rv_lv_ratio_gte_1 &lt;= 0.5 &amp; rv_lv_ratio_lt_1 &lt;= 0.5\"\n            while len(df_pos.query(rv_filter)) &gt; 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_max\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].max(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_max\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" &amp; {rv_col} == rv_max\")[\n                            rv_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_pos.index, label_cols[8:]] = df_pos[label_cols[8:]]\n        if \"1b\" in rule_breaks:\n            pe_filter = \" &amp; \".join([f\"{col} &lt;= 0.5\" for col in label_cols[5:8]])\n            while \"1b\" in consistency_check(df).index:\n                for col in label_cols[5:8]:\n                    df_pos.loc[df_pos.query(pe_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(pe_filter).index, col], s=0.1\n                    )\n                df.loc[df_pos.index, label_cols[5:8]] = df_pos[label_cols[5:8]].values\n        if \"1c\" in rule_breaks:\n            chronic_filter = \"chronic_pe &gt; 0.5 &amp; acute_and_chronic_pe &gt; 0.5\"\n            df_pos.loc[df_pos.query(chronic_filter).index, label_cols[3:5]] = softmax(\n                df_pos.query(chronic_filter)[label_cols[3:5]].values, axis=1\n            )\n            df.loc[df_pos.index, label_cols[3:5]] = df_pos[label_cols[3:5]]\n        if \"1d\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe &gt; 0.5 | indeterminate &gt; 0.5\"\n            while \"1d\" in consistency_check(df).index:\n                for col in label_cols[1:3]:\n                    df_pos.loc[df_pos.query(neg_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(neg_filter).index, col], s=-0.1\n                    )\n                df.loc[df_pos.index, label_cols[1:3]] = df_pos[label_cols[1:3]].values\n        if \"2a\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe &gt; 0.5 &amp; indeterminate &gt; 0.5\"\n            while len(df_neg.query(neg_filter)) &gt; 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_min\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].min(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_min\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_min\")[\n                            neg_col\n                        ].values,\n                        s=-0.1,\n                    )\n            neg_filter = \"negative_exam_for_pe &lt;= 0.5 &amp; indeterminate &lt;= 0.5\"\n            while len(df_neg.query(neg_filter)) &gt; 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_max\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].max(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_max\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" &amp; {neg_col} == neg_max\")[\n                            neg_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_neg.index, label_cols[1:3]] = df_neg[label_cols[1:3]]\n        if \"2b\" in rule_breaks:\n            while \"2b\" in consistency_check(df).index:\n                for col in label_cols[3:]:\n                    df_neg.loc[df_neg.query(f\"{col} &gt; 0.5\").index, col] = postprocess(\n                        df_neg.loc[df_neg.query(f\"{col} &gt; 0.5\").index, col], s=-0.1\n                    )\n                df.loc[df_neg.index, label_cols[3:]] = df_neg[label_cols[3:]].values\n    return df\n</code></pre>\n<p>Updated: <br>\nWe uploaded code on github (<a href=\"https://github.com/shimacos37/kaggle-rsna-2020-9th-solution)\" target=\"_blank\">https://github.com/shimacos37/kaggle-rsna-2020-9th-solution)</a>.</p>",
      "rawMarkdown": "Congratulations to all winners !\nThis competition was hard on me in many ways.\n\n# Solution Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fb7889dcd4b8229c53e2103cfb622a8e1%2FRSNA%202020%20solution%20(1).png?generation=1603760751579416&alt=media)\n\n# Preprocess\n\n- In the train data, no CT image have PE after 400th image. So, we used only images before 400th image.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fc51f745506070907901a148f396ae5d7%2Fdownload-20.png?generation=1603800720901371&alt=media)\n\n- For stage 1 training, we preprocessed image-level labels like following image.\n  - Before\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F41ffbb129012a32e592556611c4a155d%2Fdownload-21.png?generation=1603760930759265&alt=media)\n  - After\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F0a0fa0aacd34ca23e8b7a00e385d25b0%2Fdownload-22.png?generation=1603760956442464&alt=media)\n\n# Stage 1 training\n\n- We used 512 x 512 image + efficientnet-b5 and 384 x 384 image + efficientnet-b3 and using preprocessed labels.\n\n# Stage 2 training\n\n- Inference time was so severe because we used 512 x 512 image + efficientnet-b5. So, we subsampled 400 sequences to 200 sequences and used Deconvolution module.\n  - We got the same CV score when using 400 sequences.\n- We was not able to use various models in stage 1 because of resource. Therefore, we trained various models in stage 2.\n  - Input: b5-feature only, b3-feature only, b5-feature + b3-feature\n  - model: Conv1D, LSTM, GRU, Conv1D + LSTM\n  - output: 3 x 4 = 12 predictions\n\n# Stacking\n\n- We trained LGBM, Conv1D and GRU.\n  - We used only PE-exam when training pe_present_on_image by lgbm because image from negative PE doesn't affect competition metric.\n\n# Postprocess\n\n- We've implemented a heuristic post process.\n  - This postprocess increased CV and public score 0.002 (Private 0.160 -> 0.162)\n  - Main idea\n      - replace  `pe_present_on_image` with `1 - negative_exam_for_pe` when `1 - negative_exam_for_pe <= pe_present_on_image`\n      - repeat sigmoid -> logit -> logit += s -> sigmoid until satisfying label consistency\n```python\nlabel_cols = [\n      \"pe_present_on_image\",\n      \"negative_exam_for_pe\",\n      \"indeterminate\",\n      \"chronic_pe\",\n      \"acute_and_chronic_pe\",\n      \"central_pe\",\n      \"leftsided_pe\",\n      \"rightsided_pe\",\n      \"rv_lv_ratio_gte_1\",\n      \"rv_lv_ratio_lt_1\",\n    ]\n\ndef postprocess(x, s=2.0):\n    logit = np.log(x/(1 - x))\n    logit = logit + s\n    sigmoid = 1 / (1 + np.exp(-logit))\n    return sigmoid\n\ndef satisfy_label_consistency(df):\n    rule_breaks = consistency_check(df).index\n    print(rule_breaks)\n    if len(rule_breaks) > 0:\n        df[\"positive_exam_for_pe\"] = 1 - df[\"negative_exam_for_pe\"]\n        df.loc[\n            df.query(\"positive_exam_for_pe <= pe_present_on_image\").index,\n            \"pe_present_on_image\",\n        ] = df.loc[\n            df.query(\"positive_exam_for_pe <= pe_present_on_image\").index,\n            \"positive_exam_for_pe\",\n        ]\n        rule_breaks = consistency_check(df).index\n        df[\"positive_images_in_exam\"] = df[\"StudyInstanceUID\"].map(\n            df.groupby([\"StudyInstanceUID\"])[\"pe_present_on_image\"].max()\n        )\n        df_pos = df.query(\"positive_images_in_exam > 0.5\")\n        df_neg = df.query(\"positive_images_in_exam <= 0.5\")\n        if \"1a\" in rule_breaks:\n            rv_filter = \"rv_lv_ratio_gte_1 > 0.5 & rv_lv_ratio_lt_1 > 0.5\"\n            while len(df_pos.query(rv_filter)) > 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_min\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].min(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_min\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_min\")[\n                            rv_col\n                        ].values,\n                        s=-0.1,\n                    )\n            rv_filter = \"rv_lv_ratio_gte_1 <= 0.5 & rv_lv_ratio_lt_1 <= 0.5\"\n            while len(df_pos.query(rv_filter)) > 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_max\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].max(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_max\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_max\")[\n                            rv_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_pos.index, label_cols[8:]] = df_pos[label_cols[8:]]\n        if \"1b\" in rule_breaks:\n            pe_filter = \" & \".join([f\"{col} <= 0.5\" for col in label_cols[5:8]])\n            while \"1b\" in consistency_check(df).index:\n                for col in label_cols[5:8]:\n                    df_pos.loc[df_pos.query(pe_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(pe_filter).index, col], s=0.1\n                    )\n                df.loc[df_pos.index, label_cols[5:8]] = df_pos[label_cols[5:8]].values\n        if \"1c\" in rule_breaks:\n            chronic_filter = \"chronic_pe > 0.5 & acute_and_chronic_pe > 0.5\"\n            df_pos.loc[df_pos.query(chronic_filter).index, label_cols[3:5]] = softmax(\n                df_pos.query(chronic_filter)[label_cols[3:5]].values, axis=1\n            )\n            df.loc[df_pos.index, label_cols[3:5]] = df_pos[label_cols[3:5]]\n        if \"1d\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe > 0.5 | indeterminate > 0.5\"\n            while \"1d\" in consistency_check(df).index:\n                for col in label_cols[1:3]:\n                    df_pos.loc[df_pos.query(neg_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(neg_filter).index, col], s=-0.1\n                    )\n                df.loc[df_pos.index, label_cols[1:3]] = df_pos[label_cols[1:3]].values\n        if \"2a\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe > 0.5 & indeterminate > 0.5\"\n            while len(df_neg.query(neg_filter)) > 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_min\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].min(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_min\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_min\")[\n                            neg_col\n                        ].values,\n                        s=-0.1,\n                    )\n            neg_filter = \"negative_exam_for_pe <= 0.5 & indeterminate <= 0.5\"\n            while len(df_neg.query(neg_filter)) > 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_max\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].max(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_max\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_max\")[\n                            neg_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_neg.index, label_cols[1:3]] = df_neg[label_cols[1:3]]\n        if \"2b\" in rule_breaks:\n            while \"2b\" in consistency_check(df).index:\n                for col in label_cols[3:]:\n                    df_neg.loc[df_neg.query(f\"{col} > 0.5\").index, col] = postprocess(\n                        df_neg.loc[df_neg.query(f\"{col} > 0.5\").index, col], s=-0.1\n                    )\n                df.loc[df_neg.index, label_cols[3:]] = df_neg[label_cols[3:]].values\n    return df\n\n```\n\nUpdated: \nWe uploaded code on github (https://github.com/shimacos37/kaggle-rsna-2020-9th-solution).",
      "votes": 21
    },
    {
      "id": 1061417,
      "postDate": "2020-10-27T02:16:41.523Z",
      "content": "<p>Congratulations! You got a chance to have a presentation in RSNA.</p>",
      "rawMarkdown": "Congratulations! You got a chance to have a presentation in RSNA.",
      "votes": 3
    },
    {
      "id": 1063137,
      "postDate": "2020-10-28T14:39:19.090Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> …Thanks for sharing the solution approach.</p>",
      "rawMarkdown": "Congratulations @shimacos ...Thanks for sharing the solution approach.",
      "votes": 1
    },
    {
      "id": 1061880,
      "postDate": "2020-10-27T12:05:48.880Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> this year you get a prize for top10 😁 Interesting observation on the 400th image. </p>",
      "rawMarkdown": "Congrats @shimacos this year you get a prize for top10 😁 Interesting observation on the 400th image. ",
      "votes": 1
    },
    {
      "id": 1061843,
      "postDate": "2020-10-27T11:19:09.937Z",
      "content": "<p>Comprehensive explanations. Great job +1</p>",
      "rawMarkdown": "Comprehensive explanations. Great job +1",
      "votes": 1
    },
    {
      "id": 1061407,
      "postDate": "2020-10-27T02:00:42.757Z",
      "content": "<p>Congrats on 10th place and gold medal <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> and thanks for sharing your team solution!</p>",
      "rawMarkdown": "Congrats on 10th place and gold medal @shimacos and thanks for sharing your team solution!",
      "votes": 1
    },
    {
      "id": 1062463,
      "postDate": "2020-10-27T21:12:55.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1061417,
      "author_name": "Youhan Lee",
      "author_url": "",
      "post_date": "2020-10-27T02:16:41.523000",
      "content": "<p>Congratulations! You got a chance to have a presentation in RSNA.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1063137,
      "author_name": "Gryffindor",
      "author_url": "",
      "post_date": "2020-10-28T14:39:19.090000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> …Thanks for sharing the solution approach.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061880,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2020-10-27T12:05:48.880000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> this year you get a prize for top10 😁 Interesting observation on the 400th image. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061843,
      "author_name": "Eisa",
      "author_url": "",
      "post_date": "2020-10-27T11:19:09.937000",
      "content": "<p>Comprehensive explanations. Great job +1</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061407,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T02:00:42.757000",
      "content": "<p>Congrats on 10th place and gold medal <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> and thanks for sharing your team solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062463,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T21:12:55.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061392": "Congratulations to all winners !\nThis competition was hard on me in many ways.\n\n# Solution Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fb7889dcd4b8229c53e2103cfb622a8e1%2FRSNA%202020%20solution%20(1).png?generation=1603760751579416&alt=media)\n\n# Preprocess\n\n- In the train data, no CT image have PE after 400th image. So, we used only images before 400th image.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2Fc51f745506070907901a148f396ae5d7%2Fdownload-20.png?generation=1603800720901371&alt=media)\n\n- For stage 1 training, we preprocessed image-level labels like following image.\n  - Before\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F41ffbb129012a32e592556611c4a155d%2Fdownload-21.png?generation=1603760930759265&alt=media)\n  - After\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1227363%2F0a0fa0aacd34ca23e8b7a00e385d25b0%2Fdownload-22.png?generation=1603760956442464&alt=media)\n\n# Stage 1 training\n\n- We used 512 x 512 image + efficientnet-b5 and 384 x 384 image + efficientnet-b3 and using preprocessed labels.\n\n# Stage 2 training\n\n- Inference time was so severe because we used 512 x 512 image + efficientnet-b5. So, we subsampled 400 sequences to 200 sequences and used Deconvolution module.\n  - We got the same CV score when using 400 sequences.\n- We was not able to use various models in stage 1 because of resource. Therefore, we trained various models in stage 2.\n  - Input: b5-feature only, b3-feature only, b5-feature + b3-feature\n  - model: Conv1D, LSTM, GRU, Conv1D + LSTM\n  - output: 3 x 4 = 12 predictions\n\n# Stacking\n\n- We trained LGBM, Conv1D and GRU.\n  - We used only PE-exam when training pe_present_on_image by lgbm because image from negative PE doesn't affect competition metric.\n\n# Postprocess\n\n- We've implemented a heuristic post process.\n  - This postprocess increased CV and public score 0.002 (Private 0.160 -> 0.162)\n  - Main idea\n      - replace  `pe_present_on_image` with `1 - negative_exam_for_pe` when `1 - negative_exam_for_pe <= pe_present_on_image`\n      - repeat sigmoid -> logit -> logit += s -> sigmoid until satisfying label consistency\n```python\nlabel_cols = [\n      \"pe_present_on_image\",\n      \"negative_exam_for_pe\",\n      \"indeterminate\",\n      \"chronic_pe\",\n      \"acute_and_chronic_pe\",\n      \"central_pe\",\n      \"leftsided_pe\",\n      \"rightsided_pe\",\n      \"rv_lv_ratio_gte_1\",\n      \"rv_lv_ratio_lt_1\",\n    ]\n\ndef postprocess(x, s=2.0):\n    logit = np.log(x/(1 - x))\n    logit = logit + s\n    sigmoid = 1 / (1 + np.exp(-logit))\n    return sigmoid\n\ndef satisfy_label_consistency(df):\n    rule_breaks = consistency_check(df).index\n    print(rule_breaks)\n    if len(rule_breaks) > 0:\n        df[\"positive_exam_for_pe\"] = 1 - df[\"negative_exam_for_pe\"]\n        df.loc[\n            df.query(\"positive_exam_for_pe <= pe_present_on_image\").index,\n            \"pe_present_on_image\",\n        ] = df.loc[\n            df.query(\"positive_exam_for_pe <= pe_present_on_image\").index,\n            \"positive_exam_for_pe\",\n        ]\n        rule_breaks = consistency_check(df).index\n        df[\"positive_images_in_exam\"] = df[\"StudyInstanceUID\"].map(\n            df.groupby([\"StudyInstanceUID\"])[\"pe_present_on_image\"].max()\n        )\n        df_pos = df.query(\"positive_images_in_exam > 0.5\")\n        df_neg = df.query(\"positive_images_in_exam <= 0.5\")\n        if \"1a\" in rule_breaks:\n            rv_filter = \"rv_lv_ratio_gte_1 > 0.5 & rv_lv_ratio_lt_1 > 0.5\"\n            while len(df_pos.query(rv_filter)) > 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_min\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].min(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_min\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_min\")[\n                            rv_col\n                        ].values,\n                        s=-0.1,\n                    )\n            rv_filter = \"rv_lv_ratio_gte_1 <= 0.5 & rv_lv_ratio_lt_1 <= 0.5\"\n            while len(df_pos.query(rv_filter)) > 0:\n                df_pos.loc[df_pos.query(rv_filter).index, \"rv_max\"] = df_pos.query(\n                    rv_filter\n                )[label_cols[8:]].max(1)\n                for rv_col in label_cols[8:]:\n                    df_pos.loc[\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_max\").index, rv_col\n                    ] = postprocess(\n                        df_pos.query(rv_filter + f\" & {rv_col} == rv_max\")[\n                            rv_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_pos.index, label_cols[8:]] = df_pos[label_cols[8:]]\n        if \"1b\" in rule_breaks:\n            pe_filter = \" & \".join([f\"{col} <= 0.5\" for col in label_cols[5:8]])\n            while \"1b\" in consistency_check(df).index:\n                for col in label_cols[5:8]:\n                    df_pos.loc[df_pos.query(pe_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(pe_filter).index, col], s=0.1\n                    )\n                df.loc[df_pos.index, label_cols[5:8]] = df_pos[label_cols[5:8]].values\n        if \"1c\" in rule_breaks:\n            chronic_filter = \"chronic_pe > 0.5 & acute_and_chronic_pe > 0.5\"\n            df_pos.loc[df_pos.query(chronic_filter).index, label_cols[3:5]] = softmax(\n                df_pos.query(chronic_filter)[label_cols[3:5]].values, axis=1\n            )\n            df.loc[df_pos.index, label_cols[3:5]] = df_pos[label_cols[3:5]]\n        if \"1d\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe > 0.5 | indeterminate > 0.5\"\n            while \"1d\" in consistency_check(df).index:\n                for col in label_cols[1:3]:\n                    df_pos.loc[df_pos.query(neg_filter).index, col] = postprocess(\n                        df_pos.loc[df_pos.query(neg_filter).index, col], s=-0.1\n                    )\n                df.loc[df_pos.index, label_cols[1:3]] = df_pos[label_cols[1:3]].values\n        if \"2a\" in rule_breaks:\n            neg_filter = \"negative_exam_for_pe > 0.5 & indeterminate > 0.5\"\n            while len(df_neg.query(neg_filter)) > 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_min\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].min(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_min\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_min\")[\n                            neg_col\n                        ].values,\n                        s=-0.1,\n                    )\n            neg_filter = \"negative_exam_for_pe <= 0.5 & indeterminate <= 0.5\"\n            while len(df_neg.query(neg_filter)) > 0:\n                df_neg.loc[df_neg.query(neg_filter).index, \"neg_max\"] = df_neg.query(\n                    neg_filter\n                )[label_cols[1:3]].max(1)\n                for neg_col in label_cols[1:3]:\n                    df_neg.loc[\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_max\").index,\n                        neg_col,\n                    ] = postprocess(\n                        df_neg.query(neg_filter + f\" & {neg_col} == neg_max\")[\n                            neg_col\n                        ].values,\n                        s=0.1,\n                    )\n            df.loc[df_neg.index, label_cols[1:3]] = df_neg[label_cols[1:3]]\n        if \"2b\" in rule_breaks:\n            while \"2b\" in consistency_check(df).index:\n                for col in label_cols[3:]:\n                    df_neg.loc[df_neg.query(f\"{col} > 0.5\").index, col] = postprocess(\n                        df_neg.loc[df_neg.query(f\"{col} > 0.5\").index, col], s=-0.1\n                    )\n                df.loc[df_neg.index, label_cols[3:]] = df_neg[label_cols[3:]].values\n    return df\n\n```\n\nUpdated: \nWe uploaded code on github (https://github.com/shimacos37/kaggle-rsna-2020-9th-solution).",
    "1061417": "Congratulations! You got a chance to have a presentation in RSNA.",
    "1063137": "Congratulations @shimacos ...Thanks for sharing the solution approach.",
    "1061880": "Congrats @shimacos this year you get a prize for top10 😁 Interesting observation on the 400th image. ",
    "1061843": "Comprehensive explanations. Great job +1",
    "1061407": "Congrats on 10th place and gold medal @shimacos and thanks for sharing your team solution!",
    "1062463": ""
  }
}