{
  "id": 610290,
  "title": "Reliability and usage of metadata (from train.csv vs. DICOM)",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/610290",
  "author_name": "OshiRikan",
  "post_date": "2025-10-03T04:35:58.833000",
  "votes": 0,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello, I’m currently thinking about training a model that makes use of metadata.<br>\nBy metadata, I don’t mean values extracted directly from DICOM files, but rather fields like <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code> that are available in train.csv.</p>\n<p>However, after reading other discussions and sample code, I started to wonder:</p>\n<ul>\n<li><p>The test data does not seem to include such a CSV file, which means <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code> would need to be obtained from the DICOM files instead.</p></li>\n<li><p>In the first place, the  <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code>  values in train.csv may not be fully reliable.</p></li>\n</ul>\n<p>How are you all handling and leveraging metadata in your approaches? I’d love to hear your thoughts!</p>",
  "messages": [
    {
      "id": 3297709,
      "postDate": "2025-10-03T15:39:52.653Z",
      "content": "<p>I was wondering about the meta data as well. I tried to make a simple predictor based only on sex, age and modality. In 5-fol cv I estimated a score of around 0.65 on the training data. When I submitted it, I got something like  0.55. I’m not sure how consistent the data is either. In the train set, the age is different in the train.csv file and the actual Dicom in a lot of cases.</p>",
      "rawMarkdown": "I was wondering about the meta data as well. I tried to make a simple predictor based only on sex, age and modality. In 5-fol cv I estimated a score of around 0.65 on the training data. When I submitted it, I got something like  0.55. I’m not sure how consistent the data is either. In the train set, the age is different in the train.csv file and the actual Dicom in a lot of cases.",
      "votes": 1,
      "replies": [
        {
          "id": 3297770,
          "postDate": "2025-10-03T18:38:06.427Z",
          "content": "<p>I got 0.69 CV / 0.68 LB with DICOM metadata.</p>",
          "rawMarkdown": "I got 0.69 CV / 0.68 LB with DICOM metadata.",
          "votes": 1,
          "replies": [
            {
              "id": 3297820,
              "postDate": "2025-10-03T21:21:48.967Z",
              "content": "<p>Haha,ok. Then I guess I screwed up :)</p>",
              "rawMarkdown": "Haha,ok. Then I guess I screwed up :)"
            },
            {
              "id": 3297821,
              "postDate": "2025-10-03T21:22:53Z",
              "content": "<p>And that's only age, sex and modality?</p>",
              "rawMarkdown": "And that's only age, sex and modality?"
            },
            {
              "id": 3297962,
              "postDate": "2025-10-04T06:25:38.003Z",
              "content": "<p>No, it's DICOM metadata.</p>",
              "rawMarkdown": "No, it's DICOM metadata."
            }
          ]
        },
        {
          "id": 3297935,
          "postDate": "2025-10-04T05:20:45.633Z",
          "content": "<p>Thanks for sharing your experience!<br>\nThat’s really helpful — and I’ve come to a similar conclusion.</p>\n<ul>\n<li><p>Since the test data doesn’t include a CSV file like train.csv, it seems that we would need to obtain metadata (such as PatientSex, PatientAge, and Modality) directly from the DICOM files.</p></li>\n<li><p>However, according to the <a href=\"https://www.kaggle.com/code/ryanholbrook/rsna-aneurysm-detection-demo-submission\" target=\"_blank\">demo notebook by Ryan Holbrook</a><br>\n, these attributes (PatientSex and PatientAge) don’t actually exist in the DICOM tags of this dataset.</p></li>\n</ul>\n<p>Therefore, it might not be appropriate (or even possible) to use PatientSex and PatientAge as model inputs in this competition.</p>",
          "rawMarkdown": "Thanks for sharing your experience!\nThat’s really helpful — and I’ve come to a similar conclusion.\n\n- Since the test data doesn’t include a CSV file like train.csv, it seems that we would need to obtain metadata (such as PatientSex, PatientAge, and Modality) directly from the DICOM files.\n\n- However, according to the [demo notebook by Ryan Holbrook](https://www.kaggle.com/code/ryanholbrook/rsna-aneurysm-detection-demo-submission)\n, these attributes (PatientSex and PatientAge) don’t actually exist in the DICOM tags of this dataset.\n\nTherefore, it might not be appropriate (or even possible) to use PatientSex and PatientAge as model inputs in this competition."
        }
      ]
    },
    {
      "id": 3297429,
      "postDate": "2025-10-03T06:23:20.400Z",
      "content": "<p>Hi, seems to be you must determine your target, then base of that do feature engineering and make sure what is your approach, binay classification is one approach for putting PatientSex feature as a target!</p>",
      "rawMarkdown": "Hi, seems to be you must determine your target, then base of that do feature engineering and make sure what is your approach, binay classification is one approach for putting PatientSex feature as a target!",
      "replies": [
        {
          "id": 3297928,
          "postDate": "2025-10-04T05:03:39.297Z",
          "content": "<p>Thanks for your reply! My target is aneurysm detection (presence/absence, and possibly location prediction). I’m considering PatientAge, PatientSex, and Modality as additional metadata features to support the model, rather than using them as the prediction target.</p>\n<p>Do you think using these metadata features improves performance in practice, or is it generally better to focus only on image features?</p>",
          "rawMarkdown": "Thanks for your reply! My target is aneurysm detection (presence/absence, and possibly location prediction). I’m considering PatientAge, PatientSex, and Modality as additional metadata features to support the model, rather than using them as the prediction target.\n\nDo you think using these metadata features improves performance in practice, or is it generally better to focus only on image features?",
          "replies": [
            {
              "id": 3298015,
              "postDate": "2025-10-04T10:00:10.360Z",
              "content": "<p>let's check a category : </p>\n<ol>\n<li>Image Features: These are the visual patterns, textures, and structures a model extracts directly from the medical scan (the shape of a potential aneurysm). Deep learning models, especially Convolutional Neural Networks (CNNs), are highly effective at learning these features automatically.</li>\n<li>Metadata Features: This is contextual information that is not in the image but is relevant to the case, such as the patient's age, sex, and the imaging modality used (CTA, MRA, DSA). The question is whether including these features helps or hurts the model's performance. <br>\nso for a robust models, a combined approach is often best. The image features provide the granular visual detail, while the metadata provides important clinical context that can increase robustness and generalizability. my recommendation is to use both and see each importance on the target, then making the decision would be easier !.</li>\n</ol>",
              "rawMarkdown": "let's check a category : \n1. Image Features: These are the visual patterns, textures, and structures a model extracts directly from the medical scan (the shape of a potential aneurysm). Deep learning models, especially Convolutional Neural Networks (CNNs), are highly effective at learning these features automatically.\n2. Metadata Features: This is contextual information that is not in the image but is relevant to the case, such as the patient's age, sex, and the imaging modality used (CTA, MRA, DSA). The question is whether including these features helps or hurts the model's performance. \nso for a robust models, a combined approach is often best. The image features provide the granular visual detail, while the metadata provides important clinical context that can increase robustness and generalizability. my recommendation is to use both and see each importance on the target, then making the decision would be easier !.",
              "votes": -1
            }
          ]
        }
      ]
    },
    {
      "id": 3297422,
      "postDate": "2025-10-03T05:53:46.567Z",
      "content": "<p><code>PatientSex</code> is easy to predict.</p>",
      "rawMarkdown": "`PatientSex` is easy to predict."
    },
    {
      "id": 3297408,
      "postDate": "2025-10-03T04:35:58.833Z",
      "content": "<p>Hello, I’m currently thinking about training a model that makes use of metadata.<br>\nBy metadata, I don’t mean values extracted directly from DICOM files, but rather fields like <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code> that are available in train.csv.</p>\n<p>However, after reading other discussions and sample code, I started to wonder:</p>\n<ul>\n<li><p>The test data does not seem to include such a CSV file, which means <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code> would need to be obtained from the DICOM files instead.</p></li>\n<li><p>In the first place, the  <code>PatientAge</code>, <code>PatientSex</code>, and <code>Modality</code>  values in train.csv may not be fully reliable.</p></li>\n</ul>\n<p>How are you all handling and leveraging metadata in your approaches? I’d love to hear your thoughts!</p>",
      "rawMarkdown": "Hello, I’m currently thinking about training a model that makes use of metadata.\nBy metadata, I don’t mean values extracted directly from DICOM files, but rather fields like ```PatientAge```, ```PatientSex```, and ```Modality``` that are available in train.csv.\n\nHowever, after reading other discussions and sample code, I started to wonder:\n\n- The test data does not seem to include such a CSV file, which means ```PatientAge```, ```PatientSex```, and ```Modality``` would need to be obtained from the DICOM files instead.\n\n- In the first place, the  ```PatientAge```, ```PatientSex```, and ```Modality```  values in train.csv may not be fully reliable.\n\nHow are you all handling and leveraging metadata in your approaches? I’d love to hear your thoughts!"
    }
  ],
  "comments": [
    {
      "id": 3297709,
      "author_name": "thomas rost",
      "author_url": "",
      "post_date": "2025-10-03T15:39:52.653000",
      "content": "<p>I was wondering about the meta data as well. I tried to make a simple predictor based only on sex, age and modality. In 5-fol cv I estimated a score of around 0.65 on the training data. When I submitted it, I got something like  0.55. I’m not sure how consistent the data is either. In the train set, the age is different in the train.csv file and the actual Dicom in a lot of cases.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3297770,
          "author_name": "Ogurtsov",
          "author_url": "",
          "post_date": "2025-10-03T18:38:06.427000",
          "content": "<p>I got 0.69 CV / 0.68 LB with DICOM metadata.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3297820,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-10-03T21:21:48.967000",
              "content": "<p>Haha,ok. Then I guess I screwed up :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3297821,
              "author_name": "thomas rost",
              "author_url": "",
              "post_date": "2025-10-03T21:22:53",
              "content": "<p>And that's only age, sex and modality?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3297962,
              "author_name": "Ogurtsov",
              "author_url": "",
              "post_date": "2025-10-04T06:25:38.003000",
              "content": "<p>No, it's DICOM metadata.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3297935,
          "author_name": "OshiRikan",
          "author_url": "",
          "post_date": "2025-10-04T05:20:45.633000",
          "content": "<p>Thanks for sharing your experience!<br>\nThat’s really helpful — and I’ve come to a similar conclusion.</p>\n<ul>\n<li><p>Since the test data doesn’t include a CSV file like train.csv, it seems that we would need to obtain metadata (such as PatientSex, PatientAge, and Modality) directly from the DICOM files.</p></li>\n<li><p>However, according to the <a href=\"https://www.kaggle.com/code/ryanholbrook/rsna-aneurysm-detection-demo-submission\" target=\"_blank\">demo notebook by Ryan Holbrook</a><br>\n, these attributes (PatientSex and PatientAge) don’t actually exist in the DICOM tags of this dataset.</p></li>\n</ul>\n<p>Therefore, it might not be appropriate (or even possible) to use PatientSex and PatientAge as model inputs in this competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3297429,
      "author_name": "Zahra Alipour",
      "author_url": "",
      "post_date": "2025-10-03T06:23:20.400000",
      "content": "<p>Hi, seems to be you must determine your target, then base of that do feature engineering and make sure what is your approach, binay classification is one approach for putting PatientSex feature as a target!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3297928,
          "author_name": "OshiRikan",
          "author_url": "",
          "post_date": "2025-10-04T05:03:39.297000",
          "content": "<p>Thanks for your reply! My target is aneurysm detection (presence/absence, and possibly location prediction). I’m considering PatientAge, PatientSex, and Modality as additional metadata features to support the model, rather than using them as the prediction target.</p>\n<p>Do you think using these metadata features improves performance in practice, or is it generally better to focus only on image features?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3298015,
              "author_name": "Zahra Alipour",
              "author_url": "",
              "post_date": "2025-10-04T10:00:10.360000",
              "content": "<p>let's check a category : </p>\n<ol>\n<li>Image Features: These are the visual patterns, textures, and structures a model extracts directly from the medical scan (the shape of a potential aneurysm). Deep learning models, especially Convolutional Neural Networks (CNNs), are highly effective at learning these features automatically.</li>\n<li>Metadata Features: This is contextual information that is not in the image but is relevant to the case, such as the patient's age, sex, and the imaging modality used (CTA, MRA, DSA). The question is whether including these features helps or hurts the model's performance. <br>\nso for a robust models, a combined approach is often best. The image features provide the granular visual detail, while the metadata provides important clinical context that can increase robustness and generalizability. my recommendation is to use both and see each importance on the target, then making the decision would be easier !.</li>\n</ol>",
              "votes": -1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3297422,
      "author_name": "Ogurtsov",
      "author_url": "",
      "post_date": "2025-10-03T05:53:46.567000",
      "content": "<p><code>PatientSex</code> is easy to predict.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3297709": "I was wondering about the meta data as well. I tried to make a simple predictor based only on sex, age and modality. In 5-fol cv I estimated a score of around 0.65 on the training data. When I submitted it, I got something like  0.55. I’m not sure how consistent the data is either. In the train set, the age is different in the train.csv file and the actual Dicom in a lot of cases.",
    "3297429": "Hi, seems to be you must determine your target, then base of that do feature engineering and make sure what is your approach, binay classification is one approach for putting PatientSex feature as a target!",
    "3297422": "`PatientSex` is easy to predict.",
    "3297408": "Hello, I’m currently thinking about training a model that makes use of metadata.\nBy metadata, I don’t mean values extracted directly from DICOM files, but rather fields like ```PatientAge```, ```PatientSex```, and ```Modality``` that are available in train.csv.\n\nHowever, after reading other discussions and sample code, I started to wonder:\n\n- The test data does not seem to include such a CSV file, which means ```PatientAge```, ```PatientSex```, and ```Modality``` would need to be obtained from the DICOM files instead.\n\n- In the first place, the  ```PatientAge```, ```PatientSex```, and ```Modality```  values in train.csv may not be fully reliable.\n\nHow are you all handling and leveraging metadata in your approaches? I’d love to hear your thoughts!"
  }
}