{
  "id": 188570,
  "title": "Single model for both exam and slice labels",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/188570",
  "author_name": "Jayasuriya Senthilvelan",
  "post_date": "2020-10-04T00:55:17.939000",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi, I'm new to Kaggle, so I apologize if this question or doubt has been covered somewhere else already.</p>\n<p>As far as my understanding goes, there are two types of features: exam level (Negative for PE, Indeterminate, etc.), and slice level (PE present on image).</p>\n<p>I was wondering, how is it possible to build a single model to determine all of these labels? If I use 3D CNN, I can calculate the exam level features, but I will be unable to determine which slice(s) had the PE since the network would not report this detail. I could use a 2D CNN (each slice separately), but this would neglect temporal features of PE.</p>\n<p>The only solution I can see is to use a thick slice input (pass in 5 slices at a time, with the training labels corresponding to the center slice). How did you guys deal with this issue?</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": 1036563,
      "postDate": "2020-10-04T00:55:17.940Z",
      "content": "<p>Hi, I'm new to Kaggle, so I apologize if this question or doubt has been covered somewhere else already.</p>\n<p>As far as my understanding goes, there are two types of features: exam level (Negative for PE, Indeterminate, etc.), and slice level (PE present on image).</p>\n<p>I was wondering, how is it possible to build a single model to determine all of these labels? If I use 3D CNN, I can calculate the exam level features, but I will be unable to determine which slice(s) had the PE since the network would not report this detail. I could use a 2D CNN (each slice separately), but this would neglect temporal features of PE.</p>\n<p>The only solution I can see is to use a thick slice input (pass in 5 slices at a time, with the training labels corresponding to the center slice). How did you guys deal with this issue?</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi, I'm new to Kaggle, so I apologize if this question or doubt has been covered somewhere else already.\n\nAs far as my understanding goes, there are two types of features: exam level (Negative for PE, Indeterminate, etc.), and slice level (PE present on image).\n\nI was wondering, how is it possible to build a single model to determine all of these labels? If I use 3D CNN, I can calculate the exam level features, but I will be unable to determine which slice(s) had the PE since the network would not report this detail. I could use a 2D CNN (each slice separately), but this would neglect temporal features of PE.\n\nThe only solution I can see is to use a thick slice input (pass in 5 slices at a time, with the training labels corresponding to the center slice). How did you guys deal with this issue?\n\nThanks.\n",
      "votes": 1
    },
    {
      "id": 1036649,
      "postDate": "2020-10-04T05:49:19.160Z",
      "content": "<p>You could do this with one model, though it would take some creativity. Personally, I'm working on a multi-stage approach that uses a few different models. I am guessing the effort needed to customize something creative for this competition is why relatively few submissions so far beat the sample means (score=0.434, <a href=\"https://www.kaggle.com/paulorzp/mean-baseline)\" target=\"_blank\">https://www.kaggle.com/paulorzp/mean-baseline)</a>.</p>\n<p>But your question asked how it might be possible to do this in one model. Just for fun, I'll sketch out an idea of a single CNN model:</p>\n<p>Use CNN image recognition architecture (e.g. ResNet) but instead of 3 channel images, let each channel represent a slice of the CT scan. To make the inputs consistent we would pick a fixed number of slices to use for each series. (Since the number of slices is not constant, you'll need to pick a standard number of slices, like 200, and then fill in blanks or repeats for any series with less than 200, and for series with more than 200 images, you'd have to toss some out (or skip some and assign the same prediction as an adjacent slice).</p>\n<p>This approach would then be equivalent to a multi-label image classification problem where you estimate a probability for each of the study and image labels (9+200 labels).</p>\n<p>Input: 512x512x200 -&gt; Output: 209x1 tensor </p>\n<p>With enough data, the model would learn to be internally consistent, but with only 7,000 series in our training set, you might need to add a consistency check at the end to avoid getting disqualified (<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473)\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473)</a>.</p>",
      "rawMarkdown": "You could do this with one model, though it would take some creativity. Personally, I'm working on a multi-stage approach that uses a few different models. I am guessing the effort needed to customize something creative for this competition is why relatively few submissions so far beat the sample means (score=0.434, https://www.kaggle.com/paulorzp/mean-baseline).\n\nBut your question asked how it might be possible to do this in one model. Just for fun, I'll sketch out an idea of a single CNN model:\n\nUse CNN image recognition architecture (e.g. ResNet) but instead of 3 channel images, let each channel represent a slice of the CT scan. To make the inputs consistent we would pick a fixed number of slices to use for each series. (Since the number of slices is not constant, you'll need to pick a standard number of slices, like 200, and then fill in blanks or repeats for any series with less than 200, and for series with more than 200 images, you'd have to toss some out (or skip some and assign the same prediction as an adjacent slice).\n\nThis approach would then be equivalent to a multi-label image classification problem where you estimate a probability for each of the study and image labels (9+200 labels).\n\nInput: 512x512x200 -> Output: 209x1 tensor \n\nWith enough data, the model would learn to be internally consistent, but with only 7,000 series in our training set, you might need to add a consistency check at the end to avoid getting disqualified (https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473).\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1036649,
      "author_name": "RS Turley",
      "author_url": "",
      "post_date": "2020-10-04T05:49:19.160000",
      "content": "<p>You could do this with one model, though it would take some creativity. Personally, I'm working on a multi-stage approach that uses a few different models. I am guessing the effort needed to customize something creative for this competition is why relatively few submissions so far beat the sample means (score=0.434, <a href=\"https://www.kaggle.com/paulorzp/mean-baseline)\" target=\"_blank\">https://www.kaggle.com/paulorzp/mean-baseline)</a>.</p>\n<p>But your question asked how it might be possible to do this in one model. Just for fun, I'll sketch out an idea of a single CNN model:</p>\n<p>Use CNN image recognition architecture (e.g. ResNet) but instead of 3 channel images, let each channel represent a slice of the CT scan. To make the inputs consistent we would pick a fixed number of slices to use for each series. (Since the number of slices is not constant, you'll need to pick a standard number of slices, like 200, and then fill in blanks or repeats for any series with less than 200, and for series with more than 200 images, you'd have to toss some out (or skip some and assign the same prediction as an adjacent slice).</p>\n<p>This approach would then be equivalent to a multi-label image classification problem where you estimate a probability for each of the study and image labels (9+200 labels).</p>\n<p>Input: 512x512x200 -&gt; Output: 209x1 tensor </p>\n<p>With enough data, the model would learn to be internally consistent, but with only 7,000 series in our training set, you might need to add a consistency check at the end to avoid getting disqualified (<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473)\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473)</a>.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1036563": "Hi, I'm new to Kaggle, so I apologize if this question or doubt has been covered somewhere else already.\n\nAs far as my understanding goes, there are two types of features: exam level (Negative for PE, Indeterminate, etc.), and slice level (PE present on image).\n\nI was wondering, how is it possible to build a single model to determine all of these labels? If I use 3D CNN, I can calculate the exam level features, but I will be unable to determine which slice(s) had the PE since the network would not report this detail. I could use a 2D CNN (each slice separately), but this would neglect temporal features of PE.\n\nThe only solution I can see is to use a thick slice input (pass in 5 slices at a time, with the training labels corresponding to the center slice). How did you guys deal with this issue?\n\nThanks.\n",
    "1036649": "You could do this with one model, though it would take some creativity. Personally, I'm working on a multi-stage approach that uses a few different models. I am guessing the effort needed to customize something creative for this competition is why relatively few submissions so far beat the sample means (score=0.434, https://www.kaggle.com/paulorzp/mean-baseline).\n\nBut your question asked how it might be possible to do this in one model. Just for fun, I'll sketch out an idea of a single CNN model:\n\nUse CNN image recognition architecture (e.g. ResNet) but instead of 3 channel images, let each channel represent a slice of the CT scan. To make the inputs consistent we would pick a fixed number of slices to use for each series. (Since the number of slices is not constant, you'll need to pick a standard number of slices, like 200, and then fill in blanks or repeats for any series with less than 200, and for series with more than 200 images, you'd have to toss some out (or skip some and assign the same prediction as an adjacent slice).\n\nThis approach would then be equivalent to a multi-label image classification problem where you estimate a probability for each of the study and image labels (9+200 labels).\n\nInput: 512x512x200 -> Output: 209x1 tensor \n\nWith enough data, the model would learn to be internally consistent, but with only 7,000 series in our training set, you might need to add a consistency check at the end to avoid getting disqualified (https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183473).\n"
  }
}