{
  "id": 600036,
  "title": "Dataset Update",
  "url": "/competitions/rsna-intracranial-aneurysm-detection/discussion/600036",
  "author_name": "Ryan Holbrook",
  "post_date": "2025-08-20T15:32:47.823000",
  "votes": 30,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I am in the process of posting an update to the competition dataset. It should address most of the issues identified here in the forums. Due to the size of the dataset, it will take some time to upload and process. Expect the new dataset to be live by 5:00 pm EST.</p>\n<p>As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits. The number of series removed in the test set is ~0.5%, which, I believe, is few enough not to warrant the disruption of a rescore at this point in the competition.</p>\n<p>Also note that we changed the extension of the segmentation files from <code>.nii</code> to a <code>.nii.gz</code>. The compressed format is more standard for this type of file and reduces the size of the dataset some.</p>\n<p>Finally, we have included an upgrade to the evaluation system that should reduce the overhead in the submission system significantly. We hope you observe submissions under the new dataset to complete faster than before. (The extended 12 hour limit will remain.)</p>\n<p>Thank you for your patience as we prepared this update. Please let us know if you have any questions or concerns.</p>\n<p><strong>UPDATE 4:00pm EST:</strong> The updated dataset is now live.</p>",
  "messages": [
    {
      "id": 3272252,
      "postDate": "2025-08-20T15:32:47.823Z",
      "content": "<p>Hi everyone,</p>\n<p>I am in the process of posting an update to the competition dataset. It should address most of the issues identified here in the forums. Due to the size of the dataset, it will take some time to upload and process. Expect the new dataset to be live by 5:00 pm EST.</p>\n<p>As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits. The number of series removed in the test set is ~0.5%, which, I believe, is few enough not to warrant the disruption of a rescore at this point in the competition.</p>\n<p>Also note that we changed the extension of the segmentation files from <code>.nii</code> to a <code>.nii.gz</code>. The compressed format is more standard for this type of file and reduces the size of the dataset some.</p>\n<p>Finally, we have included an upgrade to the evaluation system that should reduce the overhead in the submission system significantly. We hope you observe submissions under the new dataset to complete faster than before. (The extended 12 hour limit will remain.)</p>\n<p>Thank you for your patience as we prepared this update. Please let us know if you have any questions or concerns.</p>\n<p><strong>UPDATE 4:00pm EST:</strong> The updated dataset is now live.</p>",
      "rawMarkdown": "Hi everyone,\n\nI am in the process of posting an update to the competition dataset. It should address most of the issues identified here in the forums. Due to the size of the dataset, it will take some time to upload and process. Expect the new dataset to be live by 5:00 pm EST.\n\nAs part of this update, we removed a small number of series with unresolvable issues in both the training and test splits. The number of series removed in the test set is ~0.5%, which, I believe, is few enough not to warrant the disruption of a rescore at this point in the competition.\n\nAlso note that we changed the extension of the segmentation files from `.nii` to a `.nii.gz`. The compressed format is more standard for this type of file and reduces the size of the dataset some.\n\nFinally, we have included an upgrade to the evaluation system that should reduce the overhead in the submission system significantly. We hope you observe submissions under the new dataset to complete faster than before. (The extended 12 hour limit will remain.)\n\nThank you for your patience as we prepared this update. Please let us know if you have any questions or concerns.\n\n**UPDATE 4:00pm EST:** The updated dataset is now live.",
      "votes": 30
    },
    {
      "id": 3272921,
      "postDate": "2025-08-21T19:00:40.680Z",
      "content": "<p>Hey all. Following up on Ryan's comment, here is what we fixed:</p>\n<h1>1. Issue with aneurysm coordinates for multi-frame DICOMS</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546</a></p>\n<ul>\n<li>Fixed with a new version train_localizer.csv with ‘f’ frame numbers for multi-frame DICOMS</li>\n</ul>\n<h1>2. Some CTs had incorrect intensity slope/intercept scaling</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636</a></p>\n<ul>\n<li>Fixed by removing a few outlier cases</li>\n<li>For some other cases, Jason Sho will provide a script to quickly fix the header variables</li>\n</ul>\n<h1>3. Some vessel segmentations were z-flipped relative to corresponding images</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857</a></p>\n<ul>\n<li>Fixed by z-flipping the affected segmentations and reuploading</li>\n</ul>\n<h1>4. Three series with aneurysms but no localizers</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685</a></p>\n<ul>\n<li>Aneurysm coordinates added</li>\n</ul>\n<h1>5. One case had modality as CT but is MRI</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566</a></p>\n<ul>\n<li>Corrected this error</li>\n</ul>\n<h1>6. One exam contained only 2 MIP images</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739</a></p>\n<ul>\n<li>This study was removed</li>\n</ul>\n<p>Please note that this fix does require you to re-download a portion of the dataset. <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> is it possible to provide a list of only the files that were changed so that participants don't have to redownload the whole dataset?</p>",
      "rawMarkdown": "Hey all. Following up on Ryan's comment, here is what we fixed:\n\n# 1. Issue with aneurysm coordinates for multi-frame DICOMS\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546\n- Fixed with a new version train_localizer.csv with ‘f’ frame numbers for multi-frame DICOMS\n\n# 2. Some CTs had incorrect intensity slope/intercept scaling\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636\n- Fixed by removing a few outlier cases\n- For some other cases, Jason Sho will provide a script to quickly fix the header variables\n\n# 3. Some vessel segmentations were z-flipped relative to corresponding images\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857\n- Fixed by z-flipping the affected segmentations and reuploading\n\n# 4. Three series with aneurysms but no localizers\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685\n- Aneurysm coordinates added\n\n# 5. One case had modality as CT but is MRI\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566\n- Corrected this error\n\n# 6. One exam contained only 2 MIP images\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739\n- This study was removed\n\nPlease note that this fix does require you to re-download a portion of the dataset. @ryanholbrook is it possible to provide a list of only the files that were changed so that participants don't have to redownload the whole dataset?",
      "votes": 16,
      "replies": [
        {
          "id": 3272978,
          "postDate": "2025-08-21T22:03:13.407Z",
          "content": "<p>Thanks for all your effort in the discussions <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> and <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>. Its great when the hosts are as into competition as the competitors 😀</p>",
          "rawMarkdown": "Thanks for all your effort in the discussions @evancalabrese and @ryanholbrook. Its great when the hosts are as into competition as the competitors 😀",
          "votes": 4
        },
        {
          "id": 3273668,
          "postDate": "2025-08-23T06:11:24.337Z",
          "content": "<p>Does 'f' start at 0 or 1?</p>",
          "rawMarkdown": "Does 'f' start at 0 or 1?",
          "votes": 5,
          "replies": [
            {
              "id": 3274506,
              "postDate": "2025-08-24T23:19:17.590Z",
              "content": "<p>Dicom standard is to start at 1</p>",
              "rawMarkdown": "Dicom standard is to start at 1",
              "votes": 3
            }
          ]
        },
        {
          "id": 3273934,
          "postDate": "2025-08-23T16:41:31.553Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3274874,
          "postDate": "2025-08-25T14:49:33.473Z",
          "content": "<p>Hello everyone.</p>\n<p>I am sharing a <a href=\"https://www.kaggle.com/code/shosys/dicom-rescale-normalizer\" target=\"_blank\">notebook + helper script</a> that parses the DICOM tags for pixel intensity metadata, which has been affecting series under item #2 above. Looking forward to everyone's continued progress!</p>",
          "rawMarkdown": "Hello everyone.\n\nI am sharing a [notebook + helper script](https://www.kaggle.com/code/shosys/dicom-rescale-normalizer) that parses the DICOM tags for pixel intensity metadata, which has been affecting series under item #2 above. Looking forward to everyone's continued progress!",
          "votes": 3
        }
      ]
    },
    {
      "id": 3273057,
      "postDate": "2025-08-22T04:05:21.783Z",
      "content": "<p>Hi, thanks for the update and all of the fixes!</p>\n<p>There seems to be another (minor) issue- there are two rows in the localizers dataframe which do not actually point to actual dicom files. Here is the code for reproducibility:</p>\n<pre><code> pathlib  \n pandas  pd\n\nlocalizers_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n\nseries_root = Path(\"/kaggle/input/rsna-intracranial-aneurysm-detection/series\")\n\n _,   localizers_df.iterrows():\n    sop_path = series_root / [\"SeriesInstanceUID\"] / f\"{row['SOPInstanceUID']}.dcm\"\n      sop_path.():\n        print(sop_path)\n</code></pre>\n<p>And here is the output:</p>\n<pre><code>/kaggle/input/rsna-intracranial-aneurysm-detection/series/...../......dcm\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/...../......dcm\n</code></pre>\n<p>It turns out these rows actually correspond to the same series (this is best seen by actually looking at the series). One series is just the other but with fewer slices. It seems that the reason there was an error in the localizers dataframe is that the SOP Instance UID for these two rows were accidentally switched.</p>\n<p>I would suggest removing the series with fewer slices and fixing the SOP Instance UID for the remaining series.</p>",
      "rawMarkdown": "Hi, thanks for the update and all of the fixes!\n\nThere seems to be another (minor) issue- there are two rows in the localizers dataframe which do not actually point to actual dicom files. Here is the code for reproducibility:\n\n```\nfrom pathlib import Path\nimport pandas as pd\n\nlocalizers_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n\nseries_root = Path(\"/kaggle/input/rsna-intracranial-aneurysm-detection/series\")\n\nfor _, row in localizers_df.iterrows():\n    sop_path = series_root / row[\"SeriesInstanceUID\"] / f\"{row['SOPInstanceUID']}.dcm\"\n    if not sop_path.exists():\n        print(sop_path)\n```\n\nAnd here is the output:\n\n```\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/1.2.826.0.1.3680043.8.498.11145695452143851764832708867797988068/1.2.826.0.1.3680043.8.498.11359680660692538603323710088085312565.dcm\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/1.2.826.0.1.3680043.8.498.35204126697881966597435252550544407444/1.2.826.0.1.3680043.8.498.50473067775982707701946022117324201859.dcm\n```\n\nIt turns out these rows actually correspond to the same series (this is best seen by actually looking at the series). One series is just the other but with fewer slices. It seems that the reason there was an error in the localizers dataframe is that the SOP Instance UID for these two rows were accidentally switched.\n\nI would suggest removing the series with fewer slices and fixing the SOP Instance UID for the remaining series.",
      "votes": 7,
      "replies": [
        {
          "id": 3273349,
          "postDate": "2025-08-22T15:45:59.340Z",
          "content": "<p>Thanks for pointing this out. We will investigate.</p>",
          "rawMarkdown": "Thanks for pointing this out. We will investigate.",
          "votes": 1
        },
        {
          "id": 3285952,
          "postDate": "2025-09-09T00:42:05.423Z",
          "content": "<p>Thanks. Now the question is, the coordnates match the SOP or the series?</p>\n<p>EDIT: I think it doesn't matter. The coordinates:</p>\n<blockquote>\n  <p>{'x': 265.4614134933777, 'y': 205.07006415562913}<br>\n  {'x': 265.44884105960267, 'y': 204.7547909768212}</p>\n</blockquote>",
          "rawMarkdown": "Thanks. Now the question is, the coordnates match the SOP or the series?\n\nEDIT: I think it doesn't matter. The coordinates:\n\n>{'x': 265.4614134933777, 'y': 205.07006415562913}\n{'x': 265.44884105960267, 'y': 204.7547909768212}"
        }
      ]
    },
    {
      "id": 3272727,
      "postDate": "2025-08-21T14:01:07.627Z",
      "content": "<p></p>\n<p></p>\n<p>Okay, the data should be back online now.</p>",
      "rawMarkdown": "~~Hi everyone,~~\n\n~~Our apologies, it seems the data viewer broke after the update. We are looking into it.~~\n\nOkay, the data should be back online now.",
      "votes": 3,
      "replies": [
        {
          "id": 3272878,
          "postDate": "2025-08-21T17:34:10.577Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Is the test dataset updated? because the submission completes in 1 hr and results are bad</p>",
          "rawMarkdown": "@ryanholbrook Is the test dataset updated? because the submission completes in 1 hr and results are bad",
          "votes": 1,
          "replies": [
            {
              "id": 3272916,
              "postDate": "2025-08-21T18:33:51.283Z",
              "content": "<p>I'm trying now and will let you know. Mine is already slightly &gt;1 hr.. im waiting .</p>",
              "rawMarkdown": "I'm trying now and will let you know. Mine is already slightly >1 hr.. im waiting ."
            },
            {
              "id": 3272927,
              "postDate": "2025-08-21T19:10:51.890Z",
              "content": "<p>I confirm I have the same experience as yours + a much worse position 😅 <br>\nbut if it helps you, yes, it finishes, with no errors and in around 1.5 hours - same as before and yes, the results are worse. </p>",
              "rawMarkdown": "I confirm I have the same experience as yours + a much worse position 😅 \nbut if it helps you, yes, it finishes, with no errors and in around 1.5 hours - same as before and yes, the results are worse. "
            },
            {
              "id": 3272999,
              "postDate": "2025-08-21T23:19:47.100Z",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> Yeah submission system is currently broken.</p>",
              "rawMarkdown": "@arunodhayan Yeah submission system is currently broken."
            },
            {
              "id": 3273104,
              "postDate": "2025-08-22T06:07:04.283Z",
              "content": "<p>Yes still it is broken</p>",
              "rawMarkdown": "Yes still it is broken"
            }
          ]
        }
      ]
    },
    {
      "id": 3272554,
      "postDate": "2025-08-21T07:57:22.863Z",
      "content": "<p>Hello team <br>\nthanks for this update. </p>\n<p>The dataset is still not available </p>\n<p>Existing notebooks - Dataset does not show. <br>\nAlso Tried creating new notebook - Fails to show Sataset </p>\n<p>I also tried adding it manually \"Add Dataset\" &gt; Competition Dataset &gt; RSNA <br>\nIt understands that needs to remove the existing old dataset, it removes it, and when I go to Add it again, it shows \"Competition Dataset not available\"<br>\nthank you so much </p>",
      "rawMarkdown": "Hello team \nthanks for this update. \n\nThe dataset is still not available \n\nExisting notebooks - Dataset does not show. \nAlso Tried creating new notebook - Fails to show Sataset \n\nI also tried adding it manually \"Add Dataset\" > Competition Dataset > RSNA \nIt understands that needs to remove the existing old dataset, it removes it, and when I go to Add it again, it shows \"Competition Dataset not available\"\nthank you so much \n\n",
      "votes": 1
    },
    {
      "id": 3272413,
      "postDate": "2025-08-20T22:33:22.740Z",
      "content": "<p>I don't see the updated dataset in the data tab, am I mistaken or is it still being updated?</p>",
      "rawMarkdown": "I don't see the updated dataset in the data tab, am I mistaken or is it still being updated?",
      "votes": 1,
      "replies": [
        {
          "id": 3272418,
          "postDate": "2025-08-20T22:53:58.650Z",
          "content": "<p>I doesn't show but it can be accessed by old original paths.</p>",
          "rawMarkdown": "I doesn't show but it can be accessed by old original paths."
        },
        {
          "id": 3272747,
          "postDate": "2025-08-21T14:35:10.950Z",
          "content": "<p>It is available in the data tab now</p>",
          "rawMarkdown": "It is available in the data tab now"
        }
      ]
    },
    {
      "id": 3272568,
      "postDate": "2025-08-21T08:33:29.177Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> On the data page, I can’t see or download the dataset.</p>",
      "rawMarkdown": "@ryanholbrook On the data page, I can’t see or download the dataset.",
      "votes": 2
    },
    {
      "id": 3272581,
      "postDate": "2025-08-21T08:57:53.447Z",
      "content": "<p>Do we have a subset of the updated data? Downloading the whole dataset took days.</p>",
      "rawMarkdown": "Do we have a subset of the updated data? Downloading the whole dataset took days.",
      "votes": 1
    },
    {
      "id": 3286831,
      "postDate": "2025-09-10T13:25:02.193Z",
      "content": "<p>The dataset selected for final scoring—is it the current one, or has new data been added?</p>",
      "rawMarkdown": "The dataset selected for final scoring—is it the current one, or has new data been added?",
      "replies": [
        {
          "id": 3287058,
          "postDate": "2025-09-10T20:00:31.100Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/liuarmy\" target=\"_blank\">@liuarmy</a> There was a post that advertised all the changes made. This one <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600036#3272921\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600036#3272921</a></p>",
          "rawMarkdown": "Hi @liuarmy There was a post that advertised all the changes made. This one \nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600036#3272921"
        }
      ]
    },
    {
      "id": 3277474,
      "postDate": "2025-08-28T02:54:59.810Z",
      "content": "<p>if i have downloaded the entire dataset before, then now i need to download it all again, right?</p>",
      "rawMarkdown": "if i have downloaded the entire dataset before, then now i need to download it all again, right?",
      "replies": [
        {
          "id": 3278020,
          "postDate": "2025-08-29T05:09:38.473Z",
          "content": "<p>I'm working on the Kaggle notebook without downloading anything, but I think you'll probably only need to download <code>train.csv</code> and <code>train_localizers.csv</code>.</p>",
          "rawMarkdown": "I'm working on the Kaggle notebook without downloading anything, but I think you'll probably only need to download `train.csv` and `train_localizers.csv`."
        }
      ]
    },
    {
      "id": 3275148,
      "postDate": "2025-08-26T02:19:03.357Z",
      "content": "<p>About <code>train.csv</code> and <code>train_localizers.csv</code> after update</p>\n<ol>\n<li><p>The <code>train_localizers.csv</code> data corresponding to the following two series that the aneurysm flag is set in <code>train.csv</code> has been deleted:</p>\n<pre><code>.....\n.....\n</code></pre>\n<ul>\n<li><p>the deleted data</p>\n<p>1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,<br>\n1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,<br>\n\"{'x': 248.09653846153847, 'y': 208.96615384615387}\",<br>\nRight Supraclinoid Internal Carotid Artery</p>\n<p>1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,<br>\n1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,<br>\n\"{'x': 248.90884615384618, 'y': 207.34153846153848}\",<br>\nRight Supraclinoid Internal Carotid Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 127.89374302532185, 'y': 134.49038577281524}\",<br>\nRight Middle Cerebral Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 129.88038461538463, 'y': 135.08307692307693}\",<br>\nRight Middle Cerebral Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 130.54500000000002, 'y': 135.74769230769232}\",<br>\nRight Middle Cerebral Artery</p></li></ul></li>\n<li><p>The <code>train_localizers.csv</code> contains data where IDs and coordinates are the same, but location is different.</p>\n<pre><code>.....,\n.....,\n,\nLeft Supraclinoid Internal Carotid Artery\n\n.....,  \n.....,  \n,  \nLeft Middle Cerebral Artery\n</code></pre>\n<ul>\n<li>The x,y coordinates of the latter were <code>(310.24964936886397, 191.73071528751754)</code> before the update.</li></ul></li>\n</ol>",
      "rawMarkdown": "About `train.csv` and `train_localizers.csv` after update\n\n1. The `train_localizers.csv` data corresponding to the following two series that the aneurysm flag is set in `train.csv` has been deleted:\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341\n  - the deleted data\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,\n        1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,\n        \"{'x': 248.09653846153847, 'y': 208.96615384615387}\",\n        Right Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,\n        1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,\n        \"{'x': 248.90884615384618, 'y': 207.34153846153848}\",\n        Right Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 127.89374302532185, 'y': 134.49038577281524}\",\n        Right Middle Cerebral Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 129.88038461538463, 'y': 135.08307692307693}\",\n        Right Middle Cerebral Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 130.54500000000002, 'y': 135.74769230769232}\",\n        Right Middle Cerebral Artery\n\n2. The `train_localizers.csv` contains data where IDs and coordinates are the same, but location is different.\n\n        1.2.826.0.1.3680043.8.498.10733938921373716882398209756836684843,\n        1.2.826.0.1.3680043.8.498.97121094041712829608625148100059963709,\n        \"{'x': 277.93548387096774, 'y': 190.2945301542777, 'f': 71}\",\n        Left Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.10733938921373716882398209756836684843,  \n        1.2.826.0.1.3680043.8.498.97121094041712829608625148100059963709,  \n        \"{'x': 277.93548387096774, 'y': 190.2945301542777, 'f': 71}\",  \n        Left Middle Cerebral Artery\n\n  - The x,y coordinates of the latter were `(310.24964936886397, 191.73071528751754)` before the update."
    },
    {
      "id": 3272487,
      "postDate": "2025-08-21T04:31:11.057Z",
      "content": "<p>I'm new to the kaggle competition. I tried to search. It didn't work. Is the data file ready? <br>\n<code>kaggle competitions list -s \"rsna\"</code></p>",
      "rawMarkdown": "I'm new to the kaggle competition. I tried to search. It didn't work. Is the data file ready? \n`kaggle competitions list -s \"rsna\"`",
      "replies": [
        {
          "id": 3272746,
          "postDate": "2025-08-21T14:34:57.187Z",
          "content": "<p>The data is ready now.</p>",
          "rawMarkdown": "The data is ready now."
        }
      ]
    },
    {
      "id": 3272448,
      "postDate": "2025-08-21T01:14:09.060Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I submit 11hr ago, and it raise \"submission file with incorrect format\". Is this happened due to data update? </p>\n<p>Update: all submit turns out to be fail prediction, go to exception prediction with all 0.1 probabilities.<br>\nAll the submission fails.</p>",
      "rawMarkdown": "@ryanholbrook I submit 11hr ago, and it raise \"submission file with incorrect format\". Is this happened due to data update? \n\nUpdate: all submit turns out to be fail prediction, go to exception prediction with all 0.1 probabilities.\nAll the submission fails."
    },
    {
      "id": 3272420,
      "postDate": "2025-08-20T23:01:22.417Z",
      "content": "<blockquote>\n  <p>As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits.</p>\n</blockquote>\n<p>I've noticed I'm pretty sure exactly double segmentations available. Are they new or just flips from original ones? Or may be are just all at one same common folder?</p>\n<p>EDIT:</p>\n<blockquote>\n  <p>['1.2.826.0.1.3680043.8.498.62169558538817009391695314359016512306.nii',<br>\n   '1.2.826.0.1.3680043.8.498.15111820005882064793593034423469604305_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.17415277997649872560329721717694101082_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.56479623144539472445940519727300319231_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.79221197357738210862579456170058377494.nii',<br>\n   '1.2.826.0.1.3680043.8.498.24941924992372724575490063788348447936.nii']</p>\n</blockquote>\n<p>Volumes and labels at same folder.</p>",
      "rawMarkdown": ">As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits.\n\nI've noticed I'm pretty sure exactly double segmentations available. Are they new or just flips from original ones? Or may be are just all at one same common folder?\n\nEDIT:\n\n>['1.2.826.0.1.3680043.8.498.62169558538817009391695314359016512306.nii',\n '1.2.826.0.1.3680043.8.498.15111820005882064793593034423469604305_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.17415277997649872560329721717694101082_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.56479623144539472445940519727300319231_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.79221197357738210862579456170058377494.nii',\n '1.2.826.0.1.3680043.8.498.24941924992372724575490063788348447936.nii']\n\nVolumes and labels at same folder."
    },
    {
      "id": 3272361,
      "postDate": "2025-08-20T19:19:53.617Z",
      "content": "<p>Is it ready ? </p>",
      "rawMarkdown": "Is it ready ? ",
      "replies": [
        {
          "id": 3272745,
          "postDate": "2025-08-21T14:34:32.533Z",
          "content": "<p>It is ready now.</p>",
          "rawMarkdown": "It is ready now.",
          "replies": [
            {
              "id": 3273336,
              "postDate": "2025-08-22T15:21:26.827Z",
              "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> Testing happened only for 1 hour and gave worse results. Was the test dataset updated?”</p>",
              "rawMarkdown": "@evancalabrese Testing happened only for 1 hour and gave worse results. Was the test dataset updated?”"
            },
            {
              "id": 3273612,
              "postDate": "2025-08-23T02:24:30.043Z",
              "content": "<p>Similar to <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>, I am having issues with submissions. </p>\n<p>Unlike Arunodhayan, my submissions do not succeed, and instead fail after 5 to 6 minutes. Local submissions (on the toy test dataset of 3 samples) succeed. </p>\n<p>Moreover, I made a submission with an old model/submission notebook that previously succeeded (7 days ago). Retrying this old submission now fails too.</p>",
              "rawMarkdown": "Similar to @arunodhayan, I am having issues with submissions. \n\nUnlike Arunodhayan, my submissions do not succeed, and instead fail after 5 to 6 minutes. Local submissions (on the toy test dataset of 3 samples) succeed. \n\nMoreover, I made a submission with an old model/submission notebook that previously succeeded (7 days ago). Retrying this old submission now fails too."
            },
            {
              "id": 3273625,
              "postDate": "2025-08-23T02:58:58.787Z",
              "content": "<p><a href=\"https://www.kaggle.com/danphan\" target=\"_blank\">@danphan</a> check this discussion <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600396\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600396</a></p>",
              "rawMarkdown": "@danphan check this discussion https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600396"
            }
          ]
        }
      ]
    },
    {
      "id": 3273303,
      "postDate": "2025-08-22T14:20:28.807Z",
      "content": "<p>Thanks for the update…!</p>",
      "rawMarkdown": "Thanks for the update...!"
    }
  ],
  "comments": [
    {
      "id": 3272921,
      "author_name": "Evan Calabrese",
      "author_url": "",
      "post_date": "2025-08-21T19:00:40.680000",
      "content": "<p>Hey all. Following up on Ryan's comment, here is what we fixed:</p>\n<h1>1. Issue with aneurysm coordinates for multi-frame DICOMS</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546</a></p>\n<ul>\n<li>Fixed with a new version train_localizer.csv with ‘f’ frame numbers for multi-frame DICOMS</li>\n</ul>\n<h1>2. Some CTs had incorrect intensity slope/intercept scaling</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636</a></p>\n<ul>\n<li>Fixed by removing a few outlier cases</li>\n<li>For some other cases, Jason Sho will provide a script to quickly fix the header variables</li>\n</ul>\n<h1>3. Some vessel segmentations were z-flipped relative to corresponding images</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857</a></p>\n<ul>\n<li>Fixed by z-flipping the affected segmentations and reuploading</li>\n</ul>\n<h1>4. Three series with aneurysms but no localizers</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685</a></p>\n<ul>\n<li>Aneurysm coordinates added</li>\n</ul>\n<h1>5. One case had modality as CT but is MRI</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566</a></p>\n<ul>\n<li>Corrected this error</li>\n</ul>\n<h1>6. One exam contained only 2 MIP images</h1>\n<p><a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739</a></p>\n<ul>\n<li>This study was removed</li>\n</ul>\n<p>Please note that this fix does require you to re-download a portion of the dataset. <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> is it possible to provide a list of only the files that were changed so that participants don't have to redownload the whole dataset?</p>",
      "votes": 16,
      "replies": [
        {
          "id": 3272978,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-08-21T22:03:13.407000",
          "content": "<p>Thanks for all your effort in the discussions <a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> and <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>. Its great when the hosts are as into competition as the competitors 😀</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3273668,
          "author_name": "Tony Hauptmann",
          "author_url": "",
          "post_date": "2025-08-23T06:11:24.337000",
          "content": "<p>Does 'f' start at 0 or 1?</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3274506,
              "author_name": "Evan Calabrese",
              "author_url": "",
              "post_date": "2025-08-24T23:19:17.590000",
              "content": "<p>Dicom standard is to start at 1</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3273934,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-08-23T16:41:31.553000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3274874,
          "author_name": "Jason Sho",
          "author_url": "",
          "post_date": "2025-08-25T14:49:33.473000",
          "content": "<p>Hello everyone.</p>\n<p>I am sharing a <a href=\"https://www.kaggle.com/code/shosys/dicom-rescale-normalizer\" target=\"_blank\">notebook + helper script</a> that parses the DICOM tags for pixel intensity metadata, which has been affecting series under item #2 above. Looking forward to everyone's continued progress!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3273057,
      "author_name": "Dan Phan",
      "author_url": "",
      "post_date": "2025-08-22T04:05:21.783000",
      "content": "<p>Hi, thanks for the update and all of the fixes!</p>\n<p>There seems to be another (minor) issue- there are two rows in the localizers dataframe which do not actually point to actual dicom files. Here is the code for reproducibility:</p>\n<pre><code> pathlib  \n pandas  pd\n\nlocalizers_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n\nseries_root = Path(\"/kaggle/input/rsna-intracranial-aneurysm-detection/series\")\n\n _,   localizers_df.iterrows():\n    sop_path = series_root / [\"SeriesInstanceUID\"] / f\"{row['SOPInstanceUID']}.dcm\"\n      sop_path.():\n        print(sop_path)\n</code></pre>\n<p>And here is the output:</p>\n<pre><code>/kaggle/input/rsna-intracranial-aneurysm-detection/series/...../......dcm\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/...../......dcm\n</code></pre>\n<p>It turns out these rows actually correspond to the same series (this is best seen by actually looking at the series). One series is just the other but with fewer slices. It seems that the reason there was an error in the localizers dataframe is that the SOP Instance UID for these two rows were accidentally switched.</p>\n<p>I would suggest removing the series with fewer slices and fixing the SOP Instance UID for the remaining series.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3273349,
          "author_name": "Evan Calabrese",
          "author_url": "",
          "post_date": "2025-08-22T15:45:59.340000",
          "content": "<p>Thanks for pointing this out. We will investigate.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3285952,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2025-09-09T00:42:05.423000",
          "content": "<p>Thanks. Now the question is, the coordnates match the SOP or the series?</p>\n<p>EDIT: I think it doesn't matter. The coordinates:</p>\n<blockquote>\n  <p>{'x': 265.4614134933777, 'y': 205.07006415562913}<br>\n  {'x': 265.44884105960267, 'y': 204.7547909768212}</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3272727,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2025-08-21T14:01:07.627000",
      "content": "<p></p>\n<p></p>\n<p>Okay, the data should be back online now.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3272878,
          "author_name": "Arunodhayan",
          "author_url": "",
          "post_date": "2025-08-21T17:34:10.577000",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> Is the test dataset updated? because the submission completes in 1 hr and results are bad</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3272916,
              "author_name": "NikolAiD153",
              "author_url": "",
              "post_date": "2025-08-21T18:33:51.283000",
              "content": "<p>I'm trying now and will let you know. Mine is already slightly &gt;1 hr.. im waiting .</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3272927,
              "author_name": "NikolAiD153",
              "author_url": "",
              "post_date": "2025-08-21T19:10:51.890000",
              "content": "<p>I confirm I have the same experience as yours + a much worse position 😅 <br>\nbut if it helps you, yes, it finishes, with no errors and in around 1.5 hours - same as before and yes, the results are worse. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3272999,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-08-21T23:19:47.100000",
              "content": "<p><a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a> Yeah submission system is currently broken.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3273104,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-08-22T06:07:04.283000",
              "content": "<p>Yes still it is broken</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3272554,
      "author_name": "NikolAiD153",
      "author_url": "",
      "post_date": "2025-08-21T07:57:22.863000",
      "content": "<p>Hello team <br>\nthanks for this update. </p>\n<p>The dataset is still not available </p>\n<p>Existing notebooks - Dataset does not show. <br>\nAlso Tried creating new notebook - Fails to show Sataset </p>\n<p>I also tried adding it manually \"Add Dataset\" &gt; Competition Dataset &gt; RSNA <br>\nIt understands that needs to remove the existing old dataset, it removes it, and when I go to Add it again, it shows \"Competition Dataset not available\"<br>\nthank you so much </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3272413,
      "author_name": "Dan Fu 42",
      "author_url": "",
      "post_date": "2025-08-20T22:33:22.740000",
      "content": "<p>I don't see the updated dataset in the data tab, am I mistaken or is it still being updated?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3272418,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2025-08-20T22:53:58.650000",
          "content": "<p>I doesn't show but it can be accessed by old original paths.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3272747,
          "author_name": "Evan Calabrese",
          "author_url": "",
          "post_date": "2025-08-21T14:35:10.950000",
          "content": "<p>It is available in the data tab now</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3272568,
      "author_name": "MOONMOON",
      "author_url": "",
      "post_date": "2025-08-21T08:33:29.177000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> On the data page, I can’t see or download the dataset.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3272581,
      "author_name": "Seeing Times",
      "author_url": "",
      "post_date": "2025-08-21T08:57:53.447000",
      "content": "<p>Do we have a subset of the updated data? Downloading the whole dataset took days.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3286831,
      "author_name": "liuarmy",
      "author_url": "",
      "post_date": "2025-09-10T13:25:02.193000",
      "content": "<p>The dataset selected for final scoring—is it the current one, or has new data been added?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3287058,
          "author_name": "NikolAiD153",
          "author_url": "",
          "post_date": "2025-09-10T20:00:31.100000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/liuarmy\" target=\"_blank\">@liuarmy</a> There was a post that advertised all the changes made. This one <br>\n<a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600036#3272921\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600036#3272921</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3277474,
      "author_name": "Cuog Nguyen",
      "author_url": "",
      "post_date": "2025-08-28T02:54:59.810000",
      "content": "<p>if i have downloaded the entire dataset before, then now i need to download it all again, right?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3278020,
          "author_name": "okabon",
          "author_url": "",
          "post_date": "2025-08-29T05:09:38.473000",
          "content": "<p>I'm working on the Kaggle notebook without downloading anything, but I think you'll probably only need to download <code>train.csv</code> and <code>train_localizers.csv</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3275148,
      "author_name": "okabon",
      "author_url": "",
      "post_date": "2025-08-26T02:19:03.357000",
      "content": "<p>About <code>train.csv</code> and <code>train_localizers.csv</code> after update</p>\n<ol>\n<li><p>The <code>train_localizers.csv</code> data corresponding to the following two series that the aneurysm flag is set in <code>train.csv</code> has been deleted:</p>\n<pre><code>.....\n.....\n</code></pre>\n<ul>\n<li><p>the deleted data</p>\n<p>1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,<br>\n1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,<br>\n\"{'x': 248.09653846153847, 'y': 208.96615384615387}\",<br>\nRight Supraclinoid Internal Carotid Artery</p>\n<p>1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,<br>\n1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,<br>\n\"{'x': 248.90884615384618, 'y': 207.34153846153848}\",<br>\nRight Supraclinoid Internal Carotid Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 127.89374302532185, 'y': 134.49038577281524}\",<br>\nRight Middle Cerebral Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 129.88038461538463, 'y': 135.08307692307693}\",<br>\nRight Middle Cerebral Artery</p>\n<p>1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,<br>\n1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,<br>\n\"{'x': 130.54500000000002, 'y': 135.74769230769232}\",<br>\nRight Middle Cerebral Artery</p></li></ul></li>\n<li><p>The <code>train_localizers.csv</code> contains data where IDs and coordinates are the same, but location is different.</p>\n<pre><code>.....,\n.....,\n,\nLeft Supraclinoid Internal Carotid Artery\n\n.....,  \n.....,  \n,  \nLeft Middle Cerebral Artery\n</code></pre>\n<ul>\n<li>The x,y coordinates of the latter were <code>(310.24964936886397, 191.73071528751754)</code> before the update.</li></ul></li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3272487,
      "author_name": "hausen",
      "author_url": "",
      "post_date": "2025-08-21T04:31:11.057000",
      "content": "<p>I'm new to the kaggle competition. I tried to search. It didn't work. Is the data file ready? <br>\n<code>kaggle competitions list -s \"rsna\"</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3272746,
          "author_name": "Evan Calabrese",
          "author_url": "",
          "post_date": "2025-08-21T14:34:57.187000",
          "content": "<p>The data is ready now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3272448,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-08-21T01:14:09.060000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I submit 11hr ago, and it raise \"submission file with incorrect format\". Is this happened due to data update? </p>\n<p>Update: all submit turns out to be fail prediction, go to exception prediction with all 0.1 probabilities.<br>\nAll the submission fails.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3272420,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-08-20T23:01:22.417000",
      "content": "<blockquote>\n  <p>As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits.</p>\n</blockquote>\n<p>I've noticed I'm pretty sure exactly double segmentations available. Are they new or just flips from original ones? Or may be are just all at one same common folder?</p>\n<p>EDIT:</p>\n<blockquote>\n  <p>['1.2.826.0.1.3680043.8.498.62169558538817009391695314359016512306.nii',<br>\n   '1.2.826.0.1.3680043.8.498.15111820005882064793593034423469604305_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.17415277997649872560329721717694101082_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.56479623144539472445940519727300319231_cowseg.nii',<br>\n   '1.2.826.0.1.3680043.8.498.79221197357738210862579456170058377494.nii',<br>\n   '1.2.826.0.1.3680043.8.498.24941924992372724575490063788348447936.nii']</p>\n</blockquote>\n<p>Volumes and labels at same folder.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3272361,
      "author_name": "Procopio",
      "author_url": "",
      "post_date": "2025-08-20T19:19:53.617000",
      "content": "<p>Is it ready ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3272745,
          "author_name": "Evan Calabrese",
          "author_url": "",
          "post_date": "2025-08-21T14:34:32.533000",
          "content": "<p>It is ready now.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3273336,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-08-22T15:21:26.827000",
              "content": "<p><a href=\"https://www.kaggle.com/evancalabrese\" target=\"_blank\">@evancalabrese</a> Testing happened only for 1 hour and gave worse results. Was the test dataset updated?”</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3273612,
              "author_name": "Dan Phan",
              "author_url": "",
              "post_date": "2025-08-23T02:24:30.043000",
              "content": "<p>Similar to <a href=\"https://www.kaggle.com/arunodhayan\" target=\"_blank\">@arunodhayan</a>, I am having issues with submissions. </p>\n<p>Unlike Arunodhayan, my submissions do not succeed, and instead fail after 5 to 6 minutes. Local submissions (on the toy test dataset of 3 samples) succeed. </p>\n<p>Moreover, I made a submission with an old model/submission notebook that previously succeeded (7 days ago). Retrying this old submission now fails too.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3273625,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2025-08-23T02:58:58.787000",
              "content": "<p><a href=\"https://www.kaggle.com/danphan\" target=\"_blank\">@danphan</a> check this discussion <a href=\"https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600396\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/600396</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3273303,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-08-22T14:20:28.807000",
      "content": "<p>Thanks for the update…!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3272252": "Hi everyone,\n\nI am in the process of posting an update to the competition dataset. It should address most of the issues identified here in the forums. Due to the size of the dataset, it will take some time to upload and process. Expect the new dataset to be live by 5:00 pm EST.\n\nAs part of this update, we removed a small number of series with unresolvable issues in both the training and test splits. The number of series removed in the test set is ~0.5%, which, I believe, is few enough not to warrant the disruption of a rescore at this point in the competition.\n\nAlso note that we changed the extension of the segmentation files from `.nii` to a `.nii.gz`. The compressed format is more standard for this type of file and reduces the size of the dataset some.\n\nFinally, we have included an upgrade to the evaluation system that should reduce the overhead in the submission system significantly. We hope you observe submissions under the new dataset to complete faster than before. (The extended 12 hour limit will remain.)\n\nThank you for your patience as we prepared this update. Please let us know if you have any questions or concerns.\n\n**UPDATE 4:00pm EST:** The updated dataset is now live.",
    "3272921": "Hey all. Following up on Ryan's comment, here is what we fixed:\n\n# 1. Issue with aneurysm coordinates for multi-frame DICOMS\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591546\n- Fixed with a new version train_localizer.csv with ‘f’ frame numbers for multi-frame DICOMS\n\n# 2. Some CTs had incorrect intensity slope/intercept scaling\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591636\n- Fixed by removing a few outlier cases\n- For some other cases, Jason Sho will provide a script to quickly fix the header variables\n\n# 3. Some vessel segmentations were z-flipped relative to corresponding images\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593857\n- Fixed by z-flipping the affected segmentations and reuploading\n\n# 4. Three series with aneurysms but no localizers\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593685\n- Aneurysm coordinates added\n\n# 5. One case had modality as CT but is MRI\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/591566\n- Corrected this error\n\n# 6. One exam contained only 2 MIP images\nhttps://www.kaggle.com/competitions/rsna-intracranial-aneurysm-detection/discussion/593739\n- This study was removed\n\nPlease note that this fix does require you to re-download a portion of the dataset. @ryanholbrook is it possible to provide a list of only the files that were changed so that participants don't have to redownload the whole dataset?",
    "3273057": "Hi, thanks for the update and all of the fixes!\n\nThere seems to be another (minor) issue- there are two rows in the localizers dataframe which do not actually point to actual dicom files. Here is the code for reproducibility:\n\n```\nfrom pathlib import Path\nimport pandas as pd\n\nlocalizers_df = pd.read_csv(\"/kaggle/input/rsna-intracranial-aneurysm-detection/train_localizers.csv\")\n\nseries_root = Path(\"/kaggle/input/rsna-intracranial-aneurysm-detection/series\")\n\nfor _, row in localizers_df.iterrows():\n    sop_path = series_root / row[\"SeriesInstanceUID\"] / f\"{row['SOPInstanceUID']}.dcm\"\n    if not sop_path.exists():\n        print(sop_path)\n```\n\nAnd here is the output:\n\n```\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/1.2.826.0.1.3680043.8.498.11145695452143851764832708867797988068/1.2.826.0.1.3680043.8.498.11359680660692538603323710088085312565.dcm\n/kaggle/input/rsna-intracranial-aneurysm-detection/series/1.2.826.0.1.3680043.8.498.35204126697881966597435252550544407444/1.2.826.0.1.3680043.8.498.50473067775982707701946022117324201859.dcm\n```\n\nIt turns out these rows actually correspond to the same series (this is best seen by actually looking at the series). One series is just the other but with fewer slices. It seems that the reason there was an error in the localizers dataframe is that the SOP Instance UID for these two rows were accidentally switched.\n\nI would suggest removing the series with fewer slices and fixing the SOP Instance UID for the remaining series.",
    "3272727": "~~Hi everyone,~~\n\n~~Our apologies, it seems the data viewer broke after the update. We are looking into it.~~\n\nOkay, the data should be back online now.",
    "3272554": "Hello team \nthanks for this update. \n\nThe dataset is still not available \n\nExisting notebooks - Dataset does not show. \nAlso Tried creating new notebook - Fails to show Sataset \n\nI also tried adding it manually \"Add Dataset\" > Competition Dataset > RSNA \nIt understands that needs to remove the existing old dataset, it removes it, and when I go to Add it again, it shows \"Competition Dataset not available\"\nthank you so much \n\n",
    "3272413": "I don't see the updated dataset in the data tab, am I mistaken or is it still being updated?",
    "3272568": "@ryanholbrook On the data page, I can’t see or download the dataset.",
    "3272581": "Do we have a subset of the updated data? Downloading the whole dataset took days.",
    "3286831": "The dataset selected for final scoring—is it the current one, or has new data been added?",
    "3277474": "if i have downloaded the entire dataset before, then now i need to download it all again, right?",
    "3275148": "About `train.csv` and `train_localizers.csv` after update\n\n1. The `train_localizers.csv` data corresponding to the following two series that the aneurysm flag is set in `train.csv` has been deleted:\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341\n  - the deleted data\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,\n        1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,\n        \"{'x': 248.09653846153847, 'y': 208.96615384615387}\",\n        Right Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.12937082136541515013380696257898978214,\n        1.2.826.0.1.3680043.8.498.41048607004454867475111720087065861246,\n        \"{'x': 248.90884615384618, 'y': 207.34153846153848}\",\n        Right Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 127.89374302532185, 'y': 134.49038577281524}\",\n        Right Middle Cerebral Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 129.88038461538463, 'y': 135.08307692307693}\",\n        Right Middle Cerebral Artery\n\n        1.2.826.0.1.3680043.8.498.86840850085811129970747331978337342341,\n        1.2.826.0.1.3680043.8.498.31624213251003577891438927087161881149,\n        \"{'x': 130.54500000000002, 'y': 135.74769230769232}\",\n        Right Middle Cerebral Artery\n\n2. The `train_localizers.csv` contains data where IDs and coordinates are the same, but location is different.\n\n        1.2.826.0.1.3680043.8.498.10733938921373716882398209756836684843,\n        1.2.826.0.1.3680043.8.498.97121094041712829608625148100059963709,\n        \"{'x': 277.93548387096774, 'y': 190.2945301542777, 'f': 71}\",\n        Left Supraclinoid Internal Carotid Artery\n\n        1.2.826.0.1.3680043.8.498.10733938921373716882398209756836684843,  \n        1.2.826.0.1.3680043.8.498.97121094041712829608625148100059963709,  \n        \"{'x': 277.93548387096774, 'y': 190.2945301542777, 'f': 71}\",  \n        Left Middle Cerebral Artery\n\n  - The x,y coordinates of the latter were `(310.24964936886397, 191.73071528751754)` before the update.",
    "3272487": "I'm new to the kaggle competition. I tried to search. It didn't work. Is the data file ready? \n`kaggle competitions list -s \"rsna\"`",
    "3272448": "@ryanholbrook I submit 11hr ago, and it raise \"submission file with incorrect format\". Is this happened due to data update? \n\nUpdate: all submit turns out to be fail prediction, go to exception prediction with all 0.1 probabilities.\nAll the submission fails.",
    "3272420": ">As part of this update, we removed a small number of series with unresolvable issues in both the training and test splits.\n\nI've noticed I'm pretty sure exactly double segmentations available. Are they new or just flips from original ones? Or may be are just all at one same common folder?\n\nEDIT:\n\n>['1.2.826.0.1.3680043.8.498.62169558538817009391695314359016512306.nii',\n '1.2.826.0.1.3680043.8.498.15111820005882064793593034423469604305_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.17415277997649872560329721717694101082_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.56479623144539472445940519727300319231_cowseg.nii',\n '1.2.826.0.1.3680043.8.498.79221197357738210862579456170058377494.nii',\n '1.2.826.0.1.3680043.8.498.24941924992372724575490063788348447936.nii']\n\nVolumes and labels at same folder.",
    "3272361": "Is it ready ? ",
    "3273303": "Thanks for the update...!"
  }
}