{
  "id": 430686,
  "title": " Handling Large Datasets for Machine Learning on Kaggle",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/430686",
  "author_name": "Srikant Nayak",
  "post_date": "2023-08-10T18:54:21.629000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello Kaggle Community,</p>\n<p>I'm new to Kaggle and excited to participate in machine learning competitions. I've encountered a challenge with the large datasets provided in some competitions, and I'm looking for advice on how to handle them effectively. Specifically, I'm wondering about the best practices for preprocessing and training models on massive datasets.</p>\n<p>Here are a few questions I have:</p>\n<p>Data Preprocessing: How do you handle feature extraction, cleaning, and transformation on large datasets to make them suitable for training?</p>\n<p>Memory Efficiency: What techniques or libraries can I use to work with large datasets in a memory-efficient manner, especially when my local machine has limited resources?</p>\n<p>Model Selection: Are there specific machine learning algorithms or models that work better with large datasets? Should I focus on specific algorithms that scale well?</p>\n<p>Sampling Strategies: Are there any strategies to work with a smaller sample of the data during the development phase before scaling up to the full dataset?</p>\n<p>Parallel Processing: How can I leverage parallel processing or distributed computing frameworks to speed up the training process for large datasets?</p>\n<p>I appreciate any tips, best practices, or resources that you can share to help me handle large datasets effectively. Thanks for your support, and I'm looking forward to learning from experienced Kagglers!</p>\n<p>Best,<br>\nSrikant Nayak</p>",
  "messages": [
    {
      "id": 2384038,
      "postDate": "2023-08-10T18:54:21.630Z",
      "content": "<p>Hello Kaggle Community,</p>\n<p>I'm new to Kaggle and excited to participate in machine learning competitions. I've encountered a challenge with the large datasets provided in some competitions, and I'm looking for advice on how to handle them effectively. Specifically, I'm wondering about the best practices for preprocessing and training models on massive datasets.</p>\n<p>Here are a few questions I have:</p>\n<p>Data Preprocessing: How do you handle feature extraction, cleaning, and transformation on large datasets to make them suitable for training?</p>\n<p>Memory Efficiency: What techniques or libraries can I use to work with large datasets in a memory-efficient manner, especially when my local machine has limited resources?</p>\n<p>Model Selection: Are there specific machine learning algorithms or models that work better with large datasets? Should I focus on specific algorithms that scale well?</p>\n<p>Sampling Strategies: Are there any strategies to work with a smaller sample of the data during the development phase before scaling up to the full dataset?</p>\n<p>Parallel Processing: How can I leverage parallel processing or distributed computing frameworks to speed up the training process for large datasets?</p>\n<p>I appreciate any tips, best practices, or resources that you can share to help me handle large datasets effectively. Thanks for your support, and I'm looking forward to learning from experienced Kagglers!</p>\n<p>Best,<br>\nSrikant Nayak</p>",
      "rawMarkdown": "Hello Kaggle Community,\n\nI'm new to Kaggle and excited to participate in machine learning competitions. I've encountered a challenge with the large datasets provided in some competitions, and I'm looking for advice on how to handle them effectively. Specifically, I'm wondering about the best practices for preprocessing and training models on massive datasets.\n\nHere are a few questions I have:\n\nData Preprocessing: How do you handle feature extraction, cleaning, and transformation on large datasets to make them suitable for training?\n\nMemory Efficiency: What techniques or libraries can I use to work with large datasets in a memory-efficient manner, especially when my local machine has limited resources?\n\nModel Selection: Are there specific machine learning algorithms or models that work better with large datasets? Should I focus on specific algorithms that scale well?\n\nSampling Strategies: Are there any strategies to work with a smaller sample of the data during the development phase before scaling up to the full dataset?\n\nParallel Processing: How can I leverage parallel processing or distributed computing frameworks to speed up the training process for large datasets?\n\n\nI appreciate any tips, best practices, or resources that you can share to help me handle large datasets effectively. Thanks for your support, and I'm looking forward to learning from experienced Kagglers!\n\nBest,\nSrikant Nayak\n\n",
      "votes": 1
    },
    {
      "id": 2385027,
      "postDate": "2023-08-11T06:33:36.203Z",
      "content": "<p>Very ranging question - hopefully you get lots of responses.</p>\n<p>So far here are a couple of my answers.</p>\n<ol>\n<li>Use the shared PNG data set as a starter.  I am using 224x224 image size for most of my playing around.  At some point a month or so from now you might want to revisit the creation of the data set.  </li>\n<li>I think I have a decent stratified set of folds, but a pretty large difference between folds for loss and accuracy, so kind of hard to identify true improvement.</li>\n<li>I don't find the dataset to be large, actually seems like a pretty small set of patients so I would not sub sample.  </li>\n<li>I do a lot of playing around running only a single fold on a 4/5 fold set.  But again, lots of noise make decision making hard.</li>\n<li>I am doing this in tensorflow with dual GPU's and using mixed precision to get some decent training speed.</li>\n</ol>\n<p>By this point in a competition a large number of folks will have leader board scores better than a 'mean' prediction.  It looks to me that only 1 person has cracked that level of performance.  </p>",
      "rawMarkdown": "Very ranging question - hopefully you get lots of responses.\n\nSo far here are a couple of my answers.\n1.  Use the shared PNG data set as a starter.  I am using 224x224 image size for most of my playing around.  At some point a month or so from now you might want to revisit the creation of the data set.  \n2.  I think I have a decent stratified set of folds, but a pretty large difference between folds for loss and accuracy, so kind of hard to identify true improvement.\n3.  I don't find the dataset to be large, actually seems like a pretty small set of patients so I would not sub sample.  \n4.  I do a lot of playing around running only a single fold on a 4/5 fold set.  But again, lots of noise make decision making hard.\n5.  I am doing this in tensorflow with dual GPU's and using mixed precision to get some decent training speed.\n\nBy this point in a competition a large number of folks will have leader board scores better than a 'mean' prediction.  It looks to me that only 1 person has cracked that level of performance.  ",
      "votes": 2,
      "replies": [
        {
          "id": 2385139,
          "postDate": "2023-08-11T07:35:09.247Z",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> Thank you for the Response :)</p>",
          "rawMarkdown": " @pcjimmmy Thank you for the Response :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2385027,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-08-11T06:33:36.203000",
      "content": "<p>Very ranging question - hopefully you get lots of responses.</p>\n<p>So far here are a couple of my answers.</p>\n<ol>\n<li>Use the shared PNG data set as a starter.  I am using 224x224 image size for most of my playing around.  At some point a month or so from now you might want to revisit the creation of the data set.  </li>\n<li>I think I have a decent stratified set of folds, but a pretty large difference between folds for loss and accuracy, so kind of hard to identify true improvement.</li>\n<li>I don't find the dataset to be large, actually seems like a pretty small set of patients so I would not sub sample.  </li>\n<li>I do a lot of playing around running only a single fold on a 4/5 fold set.  But again, lots of noise make decision making hard.</li>\n<li>I am doing this in tensorflow with dual GPU's and using mixed precision to get some decent training speed.</li>\n</ol>\n<p>By this point in a competition a large number of folks will have leader board scores better than a 'mean' prediction.  It looks to me that only 1 person has cracked that level of performance.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2385139,
          "author_name": "Srikant Nayak",
          "author_url": "",
          "post_date": "2023-08-11T07:35:09.247000",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> Thank you for the Response :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2384038": "Hello Kaggle Community,\n\nI'm new to Kaggle and excited to participate in machine learning competitions. I've encountered a challenge with the large datasets provided in some competitions, and I'm looking for advice on how to handle them effectively. Specifically, I'm wondering about the best practices for preprocessing and training models on massive datasets.\n\nHere are a few questions I have:\n\nData Preprocessing: How do you handle feature extraction, cleaning, and transformation on large datasets to make them suitable for training?\n\nMemory Efficiency: What techniques or libraries can I use to work with large datasets in a memory-efficient manner, especially when my local machine has limited resources?\n\nModel Selection: Are there specific machine learning algorithms or models that work better with large datasets? Should I focus on specific algorithms that scale well?\n\nSampling Strategies: Are there any strategies to work with a smaller sample of the data during the development phase before scaling up to the full dataset?\n\nParallel Processing: How can I leverage parallel processing or distributed computing frameworks to speed up the training process for large datasets?\n\n\nI appreciate any tips, best practices, or resources that you can share to help me handle large datasets effectively. Thanks for your support, and I'm looking forward to learning from experienced Kagglers!\n\nBest,\nSrikant Nayak\n\n",
    "2385027": "Very ranging question - hopefully you get lots of responses.\n\nSo far here are a couple of my answers.\n1.  Use the shared PNG data set as a starter.  I am using 224x224 image size for most of my playing around.  At some point a month or so from now you might want to revisit the creation of the data set.  \n2.  I think I have a decent stratified set of folds, but a pretty large difference between folds for loss and accuracy, so kind of hard to identify true improvement.\n3.  I don't find the dataset to be large, actually seems like a pretty small set of patients so I would not sub sample.  \n4.  I do a lot of playing around running only a single fold on a 4/5 fold set.  But again, lots of noise make decision making hard.\n5.  I am doing this in tensorflow with dual GPU's and using mixed precision to get some decent training speed.\n\nBy this point in a competition a large number of folks will have leader board scores better than a 'mean' prediction.  It looks to me that only 1 person has cracked that level of performance.  "
  }
}