{
  "id": 608751,
  "title": "How to get 0,93",
  "url": "/competitions/grand-xray-slam-division-a/discussion/608751",
  "author_name": "Franklin Gois",
  "post_date": "2025-09-22T02:37:29.993000",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone, I’d like to share these notebooks to help all beginners like me who are starting out on Kaggle. I’ve learned a lot from many colleagues and would like to give back to the community.<br>\nLet’s see if anyone can beat an AUROC of 0.94!<br>\nI’d like to share 2 notebooks.</p>\n<p>1 – Convert the dataset to .npy</p>\n<p><a href=\"https://www.kaggle.com/code/franklingois/converter-dataset-npy\" target=\"_blank\">https://www.kaggle.com/code/franklingois/converter-dataset-npy</a></p>\n<p>This notebook creates <code>dataset.npy</code> (<code>numpy.memmap</code>), which is very useful for accessing many small files on disk without overhead.<br>\nThis will greatly speed up training.<br>\nUnfortunately, Kaggle has a 19 GiB storage limit, which fits roughly up to 320 × 320 resolution. I downloaded the entire dataset to my computer and converted it offline to use higher resolutions.<br>\nTo optimize space, the converter uses only 1 channel (grayscale), but you can change it to 3 (it will take roughly three times more disk space). Some models that use RGB or BGR need 3 channels. What I did was keep 1 channel to stay smaller and replicate it to three channels in the model/dataset/dataloader.</p>\n<p>2 – Complete script (already with the conversion cell)</p>\n<p><a href=\"https://www.kaggle.com/code/franklingois/how-to-get-0-93\" target=\"_blank\">https://www.kaggle.com/code/franklingois/how-to-get-0-93</a></p>\n<p>This notebook already includes the conversion cell, so just run the first cell and then you can run the K-Fold training script.<br>\nIf you use it the way I did, it will produce folds with AURC between 0.9250–0.9270, which yields an AUROC of 0.93 on the leaderboard!<br>\nThe dataset and dataloader already work with <code>.npy</code> files, and the backbone used is <code>convnext_small</code> (which expects 3 channels; the script already replicates the single channel to match).</p>\n<p>I’d like to thank everyone in the community for the tremendous learning I’ve had over the past few months.<br>\nI’m a radiologist and joined Kaggle out of curiosity—today it has become my favorite hobby!</p>\n<p>😀 Warm regards to all!</p>",
  "messages": [
    {
      "id": 3292515,
      "postDate": "2025-09-22T02:37:29.993Z",
      "content": "<p>Hello everyone, I’d like to share these notebooks to help all beginners like me who are starting out on Kaggle. I’ve learned a lot from many colleagues and would like to give back to the community.<br>\nLet’s see if anyone can beat an AUROC of 0.94!<br>\nI’d like to share 2 notebooks.</p>\n<p>1 – Convert the dataset to .npy</p>\n<p><a href=\"https://www.kaggle.com/code/franklingois/converter-dataset-npy\" target=\"_blank\">https://www.kaggle.com/code/franklingois/converter-dataset-npy</a></p>\n<p>This notebook creates <code>dataset.npy</code> (<code>numpy.memmap</code>), which is very useful for accessing many small files on disk without overhead.<br>\nThis will greatly speed up training.<br>\nUnfortunately, Kaggle has a 19 GiB storage limit, which fits roughly up to 320 × 320 resolution. I downloaded the entire dataset to my computer and converted it offline to use higher resolutions.<br>\nTo optimize space, the converter uses only 1 channel (grayscale), but you can change it to 3 (it will take roughly three times more disk space). Some models that use RGB or BGR need 3 channels. What I did was keep 1 channel to stay smaller and replicate it to three channels in the model/dataset/dataloader.</p>\n<p>2 – Complete script (already with the conversion cell)</p>\n<p><a href=\"https://www.kaggle.com/code/franklingois/how-to-get-0-93\" target=\"_blank\">https://www.kaggle.com/code/franklingois/how-to-get-0-93</a></p>\n<p>This notebook already includes the conversion cell, so just run the first cell and then you can run the K-Fold training script.<br>\nIf you use it the way I did, it will produce folds with AURC between 0.9250–0.9270, which yields an AUROC of 0.93 on the leaderboard!<br>\nThe dataset and dataloader already work with <code>.npy</code> files, and the backbone used is <code>convnext_small</code> (which expects 3 channels; the script already replicates the single channel to match).</p>\n<p>I’d like to thank everyone in the community for the tremendous learning I’ve had over the past few months.<br>\nI’m a radiologist and joined Kaggle out of curiosity—today it has become my favorite hobby!</p>\n<p>😀 Warm regards to all!</p>",
      "rawMarkdown": "Hello everyone, I’d like to share these notebooks to help all beginners like me who are starting out on Kaggle. I’ve learned a lot from many colleagues and would like to give back to the community.\nLet’s see if anyone can beat an AUROC of 0.94!\nI’d like to share 2 notebooks.\n\n1 – Convert the dataset to .npy\n\nhttps://www.kaggle.com/code/franklingois/converter-dataset-npy\n\nThis notebook creates `dataset.npy` (`numpy.memmap`), which is very useful for accessing many small files on disk without overhead.\nThis will greatly speed up training.\nUnfortunately, Kaggle has a 19 GiB storage limit, which fits roughly up to 320 × 320 resolution. I downloaded the entire dataset to my computer and converted it offline to use higher resolutions.\nTo optimize space, the converter uses only 1 channel (grayscale), but you can change it to 3 (it will take roughly three times more disk space). Some models that use RGB or BGR need 3 channels. What I did was keep 1 channel to stay smaller and replicate it to three channels in the model/dataset/dataloader.\n\n2 – Complete script (already with the conversion cell)\n\nhttps://www.kaggle.com/code/franklingois/how-to-get-0-93\n\nThis notebook already includes the conversion cell, so just run the first cell and then you can run the K-Fold training script.\nIf you use it the way I did, it will produce folds with AURC between 0.9250–0.9270, which yields an AUROC of 0.93 on the leaderboard!\nThe dataset and dataloader already work with `.npy` files, and the backbone used is `convnext_small` (which expects 3 channels; the script already replicates the single channel to match).\n\n\n\nI’d like to thank everyone in the community for the tremendous learning I’ve had over the past few months.\nI’m a radiologist and joined Kaggle out of curiosity—today it has become my favorite hobby!\n\n\n\n😀 Warm regards to all!\n",
      "votes": 6
    },
    {
      "id": 3293198,
      "postDate": "2025-09-23T08:40:57.943Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/franklingois\" target=\"_blank\">@franklingois</a> for sharing your approach and the notebooks! 🙏<br>\nIt’s always great to see participants contributing their experience and helping others in the community.<br>\nWe really appreciate the spirit of collaboration and learning you’re bringing to the competition.</p>",
      "rawMarkdown": "Thank you @franklingois for sharing your approach and the notebooks! 🙏\nIt’s always great to see participants contributing their experience and helping others in the community.\nWe really appreciate the spirit of collaboration and learning you’re bringing to the competition."
    }
  ],
  "comments": [
    {
      "id": 3293198,
      "author_name": "Guntas Dhanjal",
      "author_url": "",
      "post_date": "2025-09-23T08:40:57.943000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/franklingois\" target=\"_blank\">@franklingois</a> for sharing your approach and the notebooks! 🙏<br>\nIt’s always great to see participants contributing their experience and helping others in the community.<br>\nWe really appreciate the spirit of collaboration and learning you’re bringing to the competition.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3292515": "Hello everyone, I’d like to share these notebooks to help all beginners like me who are starting out on Kaggle. I’ve learned a lot from many colleagues and would like to give back to the community.\nLet’s see if anyone can beat an AUROC of 0.94!\nI’d like to share 2 notebooks.\n\n1 – Convert the dataset to .npy\n\nhttps://www.kaggle.com/code/franklingois/converter-dataset-npy\n\nThis notebook creates `dataset.npy` (`numpy.memmap`), which is very useful for accessing many small files on disk without overhead.\nThis will greatly speed up training.\nUnfortunately, Kaggle has a 19 GiB storage limit, which fits roughly up to 320 × 320 resolution. I downloaded the entire dataset to my computer and converted it offline to use higher resolutions.\nTo optimize space, the converter uses only 1 channel (grayscale), but you can change it to 3 (it will take roughly three times more disk space). Some models that use RGB or BGR need 3 channels. What I did was keep 1 channel to stay smaller and replicate it to three channels in the model/dataset/dataloader.\n\n2 – Complete script (already with the conversion cell)\n\nhttps://www.kaggle.com/code/franklingois/how-to-get-0-93\n\nThis notebook already includes the conversion cell, so just run the first cell and then you can run the K-Fold training script.\nIf you use it the way I did, it will produce folds with AURC between 0.9250–0.9270, which yields an AUROC of 0.93 on the leaderboard!\nThe dataset and dataloader already work with `.npy` files, and the backbone used is `convnext_small` (which expects 3 channels; the script already replicates the single channel to match).\n\n\n\nI’d like to thank everyone in the community for the tremendous learning I’ve had over the past few months.\nI’m a radiologist and joined Kaggle out of curiosity—today it has become my favorite hobby!\n\n\n\n😀 Warm regards to all!\n",
    "3293198": "Thank you @franklingois for sharing your approach and the notebooks! 🙏\nIt’s always great to see participants contributing their experience and helping others in the community.\nWe really appreciate the spirit of collaboration and learning you’re bringing to the competition."
  }
}