{
  "id": 428819,
  "title": "How to work with large dataset",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/428819",
  "author_name": "golnaz ahmadvand",
  "post_date": "2023-08-03T02:13:28.940000",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone, this is my first time working with large dataset, i cant download it in google colab because disk space is just 107 GB and its not enough, could you please tell me what should i do instead? should i change my environment?</p>",
  "messages": [
    {
      "id": 2371942,
      "postDate": "2023-08-03T11:22:43.363Z",
      "content": "<p>Checkout <strong><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">Data in PNG Format</a></strong> ,  the person has converted the dataset from dicom to png and divided it in around 10 parts. which can be downloaded individually(15 - 20 GB), I think loading 1 Dataset in colab, making a pipeline and then applying it on the whole dataset might workout</p>",
      "rawMarkdown": "Checkout **[Data in PNG Format](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427)** ,  the person has converted the dataset from dicom to png and divided it in around 10 parts. which can be downloaded individually(15 - 20 GB), I think loading 1 Dataset in colab, making a pipeline and then applying it on the whole dataset might workout",
      "votes": 3,
      "replies": [
        {
          "id": 2372251,
          "postDate": "2023-08-03T15:17:44.050Z",
          "content": "<p>thank you so much! it was so helpful🙏</p>",
          "rawMarkdown": "thank you so much! it was so helpful🙏",
          "votes": 1
        }
      ]
    },
    {
      "id": 2371215,
      "postDate": "2023-08-03T02:13:28.940Z",
      "content": "<p>Hi everyone, this is my first time working with large dataset, i cant download it in google colab because disk space is just 107 GB and its not enough, could you please tell me what should i do instead? should i change my environment?</p>",
      "rawMarkdown": "Hi everyone, this is my first time working with large dataset, i cant download it in google colab because disk space is just 107 GB and its not enough, could you please tell me what should i do instead? should i change my environment?",
      "votes": 3
    },
    {
      "id": 2377709,
      "postDate": "2023-08-07T09:38:42.720Z",
      "content": "<p>Or make it to a TFRecord dataset and load it online. Using gcs_path.</p>",
      "rawMarkdown": "Or make it to a TFRecord dataset and load it online. Using gcs_path.",
      "votes": 1
    },
    {
      "id": 2384047,
      "postDate": "2023-08-10T18:59:17.823Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true,
      "replies": [
        {
          "id": 2384794,
          "postDate": "2023-08-11T04:45:47.563Z",
          "content": "<p>Ok ChatGPT</p>",
          "rawMarkdown": "Ok ChatGPT"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2371942,
      "author_name": "AyushS9020",
      "author_url": "",
      "post_date": "2023-08-03T11:22:43.363000",
      "content": "<p>Checkout <strong><a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427\" target=\"_blank\">Data in PNG Format</a></strong> ,  the person has converted the dataset from dicom to png and divided it in around 10 parts. which can be downloaded individually(15 - 20 GB), I think loading 1 Dataset in colab, making a pipeline and then applying it on the whole dataset might workout</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2372251,
          "author_name": "golnaz ahmadvand",
          "author_url": "",
          "post_date": "2023-08-03T15:17:44.050000",
          "content": "<p>thank you so much! it was so helpful🙏</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2377709,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-08-07T09:38:42.720000",
      "content": "<p>Or make it to a TFRecord dataset and load it online. Using gcs_path.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2384047,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-10T18:59:17.823000",
      "content": "",
      "votes": -2,
      "replies": [
        {
          "id": 2384794,
          "author_name": "Shaun Comino",
          "author_url": "",
          "post_date": "2023-08-11T04:45:47.563000",
          "content": "<p>Ok ChatGPT</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2371942": "Checkout **[Data in PNG Format](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/427427)** ,  the person has converted the dataset from dicom to png and divided it in around 10 parts. which can be downloaded individually(15 - 20 GB), I think loading 1 Dataset in colab, making a pipeline and then applying it on the whole dataset might workout",
    "2371215": "Hi everyone, this is my first time working with large dataset, i cant download it in google colab because disk space is just 107 GB and its not enough, could you please tell me what should i do instead? should i change my environment?",
    "2377709": "Or make it to a TFRecord dataset and load it online. Using gcs_path.",
    "2384047": ""
  }
}