{
  "id": 609017,
  "title": "Scoring error",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609017",
  "author_name": "Alan C52",
  "post_date": "2025-09-23T06:59:12.223000",
  "votes": 0,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I'm having trouble with scoring issues. I have ran the trainset in my inference pipeline and all 1100 planets pull through, all with correct rows and columns (in my Datasets \"submission (1).csv\").</p>\n<p>I've also done it on the test set via \"submission (2).cs\". All dtypes are int64 and float64 like the sample_submission.csv.</p>\n<p>submission_df.planet_id = submission_df.planet_id.astype(int)<br>\nsubmission_df = submission_df.set_index('planet_id').loc[[i for i in CFG.ordered_planets if int(i) in list(submission_df.planet_id)]].reset_index()<br>\nsubmission_df['planet_id'] = submission_df['planet_id'].astype(int)<br>\nsubmission_df = submission_df.clip(0)<br>\nsubmission_df.to_csv('submission.csv', index = False)<br>\nsubmission_df</p>\n<p>Can anyone find where this issue is occurring from?</p>\n<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> is it possible to look in my submissions and see why?</p>\n<p><a href=\"https://www.kaggle.com/datasets/alanc52/example-submissions\" target=\"_blank\">Link to dataset</a></p>",
  "messages": [
    {
      "id": 3293539,
      "postDate": "2025-09-24T05:07:37.073Z",
      "content": "<p>If you haven't done so already, you may want to explicitly check for unexpected inf and nan values, and replace them with some defaults if any are found.  </p>\n<p>Also: some competitions are set up to expect that the rows of submission.csv are in a particular order, usually the order given in sample_submission.csv.  I've only just fired off my first submission to this one, so I'm not sure if that's a requirement here too, but it's something to check.</p>",
      "rawMarkdown": "If you haven't done so already, you may want to explicitly check for unexpected inf and nan values, and replace them with some defaults if any are found.  \n\nAlso: some competitions are set up to expect that the rows of submission.csv are in a particular order, usually the order given in sample_submission.csv.  I've only just fired off my first submission to this one, so I'm not sure if that's a requirement here too, but it's something to check.",
      "votes": 1,
      "replies": [
        {
          "id": 3293670,
          "postDate": "2025-09-24T11:43:13Z",
          "content": "<p>Thanks, yea just found out yesterday (after I made this post) I have some nan somewhere.</p>\n<p>I think I know which part of the code it comes from, but doing some tests I can't really reproduce it which sucks. And not much time left.</p>\n<p>All my training data is fine - so probably just a really far edge case.</p>",
          "rawMarkdown": "Thanks, yea just found out yesterday (after I made this post) I have some nan somewhere.\n\nI think I know which part of the code it comes from, but doing some tests I can't really reproduce it which sucks. And not much time left.\n\nAll my training data is fine - so probably just a really far edge case.",
          "replies": [
            {
              "id": 3293715,
              "postDate": "2025-09-24T13:51:31.960Z",
              "content": "<p>Fixing a nan that you can't reproduce locally is hard.  But to fix the scoring error it's causing, you can do this:</p>\n<pre><code>=np.where(np.isnan(final_means),naive_mean,final_means)\n=np.where(np.isinf(final_means),naive_mean,final_means)\n=np.where(np.isnan(final_sigmas),naive_sigma,final_sigmas)\n=np.where(np.isinf(final_sigmas),naive_sigma,final_sigmas)\n</code></pre>\n<p>Also, after you generate the submission.csv, this is a good check to do:</p>\n<pre><code> (,)  f:\n        f:\n        print()\n</code></pre>\n<p>My first submission failed, and I think it was because <code>submission.to_csv()</code> was not printing out the column header for the planet_id column.  This check will catch stuff like that, and it would have been nice if I'd had the idea to do it yesterday.  🤦 </p>",
              "rawMarkdown": "Fixing a nan that you can't reproduce locally is hard.  But to fix the scoring error it's causing, you can do this:\n\n```\nfinal_means=np.where(np.isnan(final_means),naive_mean,final_means)\nfinal_means=np.where(np.isinf(final_means),naive_mean,final_means)\nfinal_sigmas=np.where(np.isnan(final_sigmas),naive_sigma,final_sigmas)\nfinal_sigmas=np.where(np.isinf(final_sigmas),naive_sigma,final_sigmas)\n```\n\nAlso, after you generate the submission.csv, this is a good check to do:\n\n```\nwith open(\"submission.csv\",\"r\") as f:\n     for line in f:\n        print(line)\n```\n\n\nMy first submission failed, and I think it was because `submission.to_csv()` was not printing out the column header for the planet_id column.  This check will catch stuff like that, and it would have been nice if I'd had the idea to do it yesterday.  🤦 "
            },
            {
              "id": 3293803,
              "postDate": "2025-09-24T17:04:55.603Z",
              "content": "<p>Yea, this is what I resorted to and put high sigma values on those ones. I do mean of 1: 284 (if not all nan) with sizable sigma, and then mean of col if it's all nan with high sigma.</p>",
              "rawMarkdown": "Yea, this is what I resorted to and put high sigma values on those ones. I do mean of 1: 284 (if not all nan) with sizable sigma, and then mean of col if it's all nan with high sigma."
            }
          ]
        }
      ]
    },
    {
      "id": 3293465,
      "postDate": "2025-09-23T23:01:43.223Z",
      "content": "<p>Are you preventing negatives? maybe there's a weird sample in the test set that you can't see that's throwing a negative - that's the main scoring issue I had before</p>",
      "rawMarkdown": "Are you preventing negatives? maybe there's a weird sample in the test set that you can't see that's throwing a negative - that's the main scoring issue I had before"
    },
    {
      "id": 3293161,
      "postDate": "2025-09-23T06:59:12.223Z",
      "content": "<p>Hi,</p>\n<p>I'm having trouble with scoring issues. I have ran the trainset in my inference pipeline and all 1100 planets pull through, all with correct rows and columns (in my Datasets \"submission (1).csv\").</p>\n<p>I've also done it on the test set via \"submission (2).cs\". All dtypes are int64 and float64 like the sample_submission.csv.</p>\n<p>submission_df.planet_id = submission_df.planet_id.astype(int)<br>\nsubmission_df = submission_df.set_index('planet_id').loc[[i for i in CFG.ordered_planets if int(i) in list(submission_df.planet_id)]].reset_index()<br>\nsubmission_df['planet_id'] = submission_df['planet_id'].astype(int)<br>\nsubmission_df = submission_df.clip(0)<br>\nsubmission_df.to_csv('submission.csv', index = False)<br>\nsubmission_df</p>\n<p>Can anyone find where this issue is occurring from?</p>\n<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> is it possible to look in my submissions and see why?</p>\n<p><a href=\"https://www.kaggle.com/datasets/alanc52/example-submissions\" target=\"_blank\">Link to dataset</a></p>",
      "rawMarkdown": "Hi,\n\nI'm having trouble with scoring issues. I have ran the trainset in my inference pipeline and all 1100 planets pull through, all with correct rows and columns (in my Datasets \"submission (1).csv\").\n\nI've also done it on the test set via \"submission (2).cs\". All dtypes are int64 and float64 like the sample_submission.csv.\n\nsubmission_df.planet_id = submission_df.planet_id.astype(int)\nsubmission_df = submission_df.set_index('planet_id').loc[[i for i in CFG.ordered_planets if int(i) in list(submission_df.planet_id)]].reset_index()\nsubmission_df['planet_id'] = submission_df['planet_id'].astype(int)\nsubmission_df = submission_df.clip(0)\nsubmission_df.to_csv('submission.csv', index = False)\nsubmission_df\n\n\nCan anyone find where this issue is occurring from?\n\n\n@gordonyip is it possible to look in my submissions and see why?\n\n\n[Link to dataset](https://www.kaggle.com/datasets/alanc52/example-submissions)"
    }
  ],
  "comments": [
    {
      "id": 3293539,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "2025-09-24T05:07:37.073000",
      "content": "<p>If you haven't done so already, you may want to explicitly check for unexpected inf and nan values, and replace them with some defaults if any are found.  </p>\n<p>Also: some competitions are set up to expect that the rows of submission.csv are in a particular order, usually the order given in sample_submission.csv.  I've only just fired off my first submission to this one, so I'm not sure if that's a requirement here too, but it's something to check.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3293670,
          "author_name": "Alan C52",
          "author_url": "",
          "post_date": "2025-09-24T11:43:13",
          "content": "<p>Thanks, yea just found out yesterday (after I made this post) I have some nan somewhere.</p>\n<p>I think I know which part of the code it comes from, but doing some tests I can't really reproduce it which sucks. And not much time left.</p>\n<p>All my training data is fine - so probably just a really far edge case.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3293715,
              "author_name": "particlebbq",
              "author_url": "",
              "post_date": "2025-09-24T13:51:31.960000",
              "content": "<p>Fixing a nan that you can't reproduce locally is hard.  But to fix the scoring error it's causing, you can do this:</p>\n<pre><code>=np.where(np.isnan(final_means),naive_mean,final_means)\n=np.where(np.isinf(final_means),naive_mean,final_means)\n=np.where(np.isnan(final_sigmas),naive_sigma,final_sigmas)\n=np.where(np.isinf(final_sigmas),naive_sigma,final_sigmas)\n</code></pre>\n<p>Also, after you generate the submission.csv, this is a good check to do:</p>\n<pre><code> (,)  f:\n        f:\n        print()\n</code></pre>\n<p>My first submission failed, and I think it was because <code>submission.to_csv()</code> was not printing out the column header for the planet_id column.  This check will catch stuff like that, and it would have been nice if I'd had the idea to do it yesterday.  🤦 </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3293803,
              "author_name": "Alan C52",
              "author_url": "",
              "post_date": "2025-09-24T17:04:55.603000",
              "content": "<p>Yea, this is what I resorted to and put high sigma values on those ones. I do mean of 1: 284 (if not all nan) with sizable sigma, and then mean of col if it's all nan with high sigma.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3293465,
      "author_name": "D. DiMonte",
      "author_url": "",
      "post_date": "2025-09-23T23:01:43.223000",
      "content": "<p>Are you preventing negatives? maybe there's a weird sample in the test set that you can't see that's throwing a negative - that's the main scoring issue I had before</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3293539": "If you haven't done so already, you may want to explicitly check for unexpected inf and nan values, and replace them with some defaults if any are found.  \n\nAlso: some competitions are set up to expect that the rows of submission.csv are in a particular order, usually the order given in sample_submission.csv.  I've only just fired off my first submission to this one, so I'm not sure if that's a requirement here too, but it's something to check.",
    "3293465": "Are you preventing negatives? maybe there's a weird sample in the test set that you can't see that's throwing a negative - that's the main scoring issue I had before",
    "3293161": "Hi,\n\nI'm having trouble with scoring issues. I have ran the trainset in my inference pipeline and all 1100 planets pull through, all with correct rows and columns (in my Datasets \"submission (1).csv\").\n\nI've also done it on the test set via \"submission (2).cs\". All dtypes are int64 and float64 like the sample_submission.csv.\n\nsubmission_df.planet_id = submission_df.planet_id.astype(int)\nsubmission_df = submission_df.set_index('planet_id').loc[[i for i in CFG.ordered_planets if int(i) in list(submission_df.planet_id)]].reset_index()\nsubmission_df['planet_id'] = submission_df['planet_id'].astype(int)\nsubmission_df = submission_df.clip(0)\nsubmission_df.to_csv('submission.csv', index = False)\nsubmission_df\n\n\nCan anyone find where this issue is occurring from?\n\n\n@gordonyip is it possible to look in my submissions and see why?\n\n\n[Link to dataset](https://www.kaggle.com/datasets/alanc52/example-submissions)"
  }
}