{
  "id": 216570,
  "title": "Extracting DICOM non-image metadata",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/216570",
  "author_name": "Phan Nguyen",
  "post_date": "2021-02-03T09:13:31.803000",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is there anyone willing to extract the DICOM files' metadata? I'm having trouble getting a baseline solution to work because the files are too big. I am already using the image data extracted, so having this extra data would be very helpful.</p>\n<p>Thanks. </p>",
  "messages": [
    {
      "id": 1184076,
      "postDate": "2021-02-03T11:30:26.233Z",
      "content": "<p>Have a look at <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/output#6.-What-is-in-the-.dicom-meta-data?\" target=\"_blank\">this notebook section</a>, where I did that (I believe there's quite a few others that have done this and may have other approaches - look around for what looks good to you). In fact, the notebook saves the results as <code>train_dicom_properties.csv.bz2</code> and this is a notebook output. So, you can add this notebook as data to your own notebook on Kaggle (see Add Data button in notebook) and you should just be able to read that with pandas <code>pd.read_csv</code>.</p>",
      "rawMarkdown": "Have a look at [this notebook section](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/output#6.-What-is-in-the-.dicom-meta-data?), where I did that (I believe there's quite a few others that have done this and may have other approaches - look around for what looks good to you). In fact, the notebook saves the results as `train_dicom_properties.csv.bz2` and this is a notebook output. So, you can add this notebook as data to your own notebook on Kaggle (see Add Data button in notebook) and you should just be able to read that with pandas `pd.read_csv`.",
      "votes": 1,
      "replies": [
        {
          "id": 1186184,
          "postDate": "2021-02-04T16:18:10.963Z",
          "content": "<p>Thanks a lot, I'll look into it. </p>",
          "rawMarkdown": "Thanks a lot, I'll look into it. "
        }
      ]
    },
    {
      "id": 1185191,
      "postDate": "2021-02-04T02:58:06.677Z",
      "content": "<p>you want standard data, not resized ?</p>",
      "rawMarkdown": "you want standard data, not resized ?",
      "votes": 2,
      "replies": [
        {
          "id": 1185727,
          "postDate": "2021-02-04T10:35:26.227Z",
          "content": "<p>Well the other metadata shouldn't be too large, and I'm not talking about the image data yes just the standard data. </p>",
          "rawMarkdown": "Well the other metadata shouldn't be too large, and I'm not talking about the image data yes just the standard data. "
        },
        {
          "id": 1185735,
          "postDate": "2021-02-04T10:40:47.947Z",
          "content": "<p>What is the maximum capacity you can use ? <a href=\"https://www.kaggle.com/phantnguyen\" target=\"_blank\">@phantnguyen</a> </p>",
          "rawMarkdown": "What is the maximum capacity you can use ? @phantnguyen "
        },
        {
          "id": 1185737,
          "postDate": "2021-02-04T10:41:55.373Z",
          "content": "<p>I can help you convert the image format <a href=\"https://www.kaggle.com/phantnguyen\" target=\"_blank\">@phantnguyen</a> </p>",
          "rawMarkdown": "I can help you convert the image format @phantnguyen "
        },
        {
          "id": 1186188,
          "postDate": "2021-02-04T16:21:14.917Z",
          "content": "<p>Thank you for the help but I already have the images sorted (I'm using the jpg converted version by another user). I just think that there may be useful information in gender or other metadata that could be included in a regression component somewhere, so that's what I want to get. I can't download the full DICOM dataset since my laptop is at near full capacity, and Colab's disk space is only 100GB, so that doesn't work either. </p>",
          "rawMarkdown": "Thank you for the help but I already have the images sorted (I'm using the jpg converted version by another user). I just think that there may be useful information in gender or other metadata that could be included in a regression component somewhere, so that's what I want to get. I can't download the full DICOM dataset since my laptop is at near full capacity, and Colab's disk space is only 100GB, so that doesn't work either. "
        }
      ]
    },
    {
      "id": 1183869,
      "postDate": "2021-02-03T09:13:31.803Z",
      "content": "<p>Is there anyone willing to extract the DICOM files' metadata? I'm having trouble getting a baseline solution to work because the files are too big. I am already using the image data extracted, so having this extra data would be very helpful.</p>\n<p>Thanks. </p>",
      "rawMarkdown": "Is there anyone willing to extract the DICOM files' metadata? I'm having trouble getting a baseline solution to work because the files are too big. I am already using the image data extracted, so having this extra data would be very helpful.\n\nThanks. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1184076,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-03T11:30:26.233000",
      "content": "<p>Have a look at <a href=\"https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/output#6.-What-is-in-the-.dicom-meta-data?\" target=\"_blank\">this notebook section</a>, where I did that (I believe there's quite a few others that have done this and may have other approaches - look around for what looks good to you). In fact, the notebook saves the results as <code>train_dicom_properties.csv.bz2</code> and this is a notebook output. So, you can add this notebook as data to your own notebook on Kaggle (see Add Data button in notebook) and you should just be able to read that with pandas <code>pd.read_csv</code>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1186184,
          "author_name": "Phan Nguyen",
          "author_url": "",
          "post_date": "2021-02-04T16:18:10.963000",
          "content": "<p>Thanks a lot, I'll look into it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1185191,
      "author_name": "Le Trong Hieu",
      "author_url": "",
      "post_date": "2021-02-04T02:58:06.677000",
      "content": "<p>you want standard data, not resized ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1185727,
          "author_name": "Phan Nguyen",
          "author_url": "",
          "post_date": "2021-02-04T10:35:26.227000",
          "content": "<p>Well the other metadata shouldn't be too large, and I'm not talking about the image data yes just the standard data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1185735,
          "author_name": "Le Trong Hieu",
          "author_url": "",
          "post_date": "2021-02-04T10:40:47.947000",
          "content": "<p>What is the maximum capacity you can use ? <a href=\"https://www.kaggle.com/phantnguyen\" target=\"_blank\">@phantnguyen</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1185737,
          "author_name": "Le Trong Hieu",
          "author_url": "",
          "post_date": "2021-02-04T10:41:55.373000",
          "content": "<p>I can help you convert the image format <a href=\"https://www.kaggle.com/phantnguyen\" target=\"_blank\">@phantnguyen</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1186188,
          "author_name": "Phan Nguyen",
          "author_url": "",
          "post_date": "2021-02-04T16:21:14.917000",
          "content": "<p>Thank you for the help but I already have the images sorted (I'm using the jpg converted version by another user). I just think that there may be useful information in gender or other metadata that could be included in a regression component somewhere, so that's what I want to get. I can't download the full DICOM dataset since my laptop is at near full capacity, and Colab's disk space is only 100GB, so that doesn't work either. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1184076": "Have a look at [this notebook section](https://www.kaggle.com/bjoernholzhauer/eda-dicom-reading-vinbigdata-chest-x-ray/output#6.-What-is-in-the-.dicom-meta-data?), where I did that (I believe there's quite a few others that have done this and may have other approaches - look around for what looks good to you). In fact, the notebook saves the results as `train_dicom_properties.csv.bz2` and this is a notebook output. So, you can add this notebook as data to your own notebook on Kaggle (see Add Data button in notebook) and you should just be able to read that with pandas `pd.read_csv`.",
    "1185191": "you want standard data, not resized ?",
    "1183869": "Is there anyone willing to extract the DICOM files' metadata? I'm having trouble getting a baseline solution to work because the files are too big. I am already using the image data extracted, so having this extra data would be very helpful.\n\nThanks. "
  }
}