{
  "id": 534389,
  "title": "submission scoring error ",
  "url": "/competitions/ariel-data-challenge-2024/discussion/534389",
  "author_name": "Makoto Kine",
  "post_date": "2024-09-16T12:04:35.253000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The submission scoring error continues.<br>\nThe code works fine on the provided single test data, and I have confirmed that there are no values below zero (both the predicted values and sigma are clipped at 0, so they won't go below zero).<br>\nI am using the following code for inference.</p>\n<p>Please let me know if you notice anything.</p>\n<pre><code> torch.utils.data  TensorDataset, DataLoader  \n\n\ndataset = \nadc_info = pd.read_csv(+, index_col=)\ntest_pre_train = np.concatenate([\n    preproc(, adc_info, , *), \n    preproc(, adc_info, , )], \n    axis=\n)\n\n\ntest_data = test_pre_train.copy()\n i  ((test_pre_train)):\n    p1, p2 = phase_detector(test_pre_train[i,:,:].mean(axis=))\n    test_data[i] = (test_data[i] - test_pre_train[i, p1:p2].mean(axis=)) / test_pre_train[i, ((p1-)) + ((p2+, ))].mean(axis=) * \n\n\ntest_x = torch.from_numpy(test_data).unsqueeze().()  \ntest_dataset = TensorDataset(test_x)\ntest_loader = DataLoader(test_dataset, batch_size=, shuffle=)\n\n\n\n\n (, )  f:\n    model = pickle.load(f)\nmodel = model.cuda()  \nmodel.()\n\n\n\npreds = np.zeros(((test_dataset), ))\nsigmas = np.zeros(((test_dataset), ))\n\n\n torch.no_grad():\n     i, data  (test_loader):\n        inputs = data[].cuda()\n        outputs = model(inputs).reshape((inputs.shape[], , ))\n\n        \n        preds[i * test_loader.batch_size: (i + ) * test_loader.batch_size] = outputs[:, :, ].detach().cpu().numpy()\n        sigmas[i * test_loader.batch_size: (i + ) * test_loader.batch_size] = (\n            outputs[:, :, ].detach().cpu().numpy() - outputs[:, :, ].detach().cpu().numpy()\n        ) \n\n\nss = pd.read_csv()\npreds = preds.clip()  \n\n\nsigmas = sigmas.clip()\n\n\nsubmission = pd.DataFrame(np.concatenate([preds, sigmas], axis=), columns=ss.columns[:])\nsubmission.index = adc_info.index\n\n\nsubmission.to_csv()\n</code></pre>",
  "messages": [
    {
      "id": 2990464,
      "postDate": "2024-09-16T12:04:35.253Z",
      "content": "<p>The submission scoring error continues.<br>\nThe code works fine on the provided single test data, and I have confirmed that there are no values below zero (both the predicted values and sigma are clipped at 0, so they won't go below zero).<br>\nI am using the following code for inference.</p>\n<p>Please let me know if you notice anything.</p>\n<pre><code> torch.utils.data  TensorDataset, DataLoader  \n\n\ndataset = \nadc_info = pd.read_csv(+, index_col=)\ntest_pre_train = np.concatenate([\n    preproc(, adc_info, , *), \n    preproc(, adc_info, , )], \n    axis=\n)\n\n\ntest_data = test_pre_train.copy()\n i  ((test_pre_train)):\n    p1, p2 = phase_detector(test_pre_train[i,:,:].mean(axis=))\n    test_data[i] = (test_data[i] - test_pre_train[i, p1:p2].mean(axis=)) / test_pre_train[i, ((p1-)) + ((p2+, ))].mean(axis=) * \n\n\ntest_x = torch.from_numpy(test_data).unsqueeze().()  \ntest_dataset = TensorDataset(test_x)\ntest_loader = DataLoader(test_dataset, batch_size=, shuffle=)\n\n\n\n\n (, )  f:\n    model = pickle.load(f)\nmodel = model.cuda()  \nmodel.()\n\n\n\npreds = np.zeros(((test_dataset), ))\nsigmas = np.zeros(((test_dataset), ))\n\n\n torch.no_grad():\n     i, data  (test_loader):\n        inputs = data[].cuda()\n        outputs = model(inputs).reshape((inputs.shape[], , ))\n\n        \n        preds[i * test_loader.batch_size: (i + ) * test_loader.batch_size] = outputs[:, :, ].detach().cpu().numpy()\n        sigmas[i * test_loader.batch_size: (i + ) * test_loader.batch_size] = (\n            outputs[:, :, ].detach().cpu().numpy() - outputs[:, :, ].detach().cpu().numpy()\n        ) \n\n\nss = pd.read_csv()\npreds = preds.clip()  \n\n\nsigmas = sigmas.clip()\n\n\nsubmission = pd.DataFrame(np.concatenate([preds, sigmas], axis=), columns=ss.columns[:])\nsubmission.index = adc_info.index\n\n\nsubmission.to_csv()\n</code></pre>",
      "rawMarkdown": "The submission scoring error continues.\nThe code works fine on the provided single test data, and I have confirmed that there are no values below zero (both the predicted values and sigma are clipped at 0, so they won't go below zero).\nI am using the following code for inference.\n\nPlease let me know if you notice anything.\n\n```python\nfrom torch.utils.data import TensorDataset, DataLoader  # Added\n\n# Loading the test data\ndataset = 'test'\nadc_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/'+f'{dataset}_adc_info.csv', index_col='planet_id')\ntest_pre_train = np.concatenate([\n    preproc(f'{dataset}', adc_info, \"FGS1\", 30*12), \n    preproc(f'{dataset}', adc_info, \"AIRS-CH0\", 30)], \n    axis=2\n)\n\n# Generate features using the phase_detector function\ntest_data = test_pre_train.copy()\nfor i in range(len(test_pre_train)):\n    p1, p2 = phase_detector(test_pre_train[i,:,1:].mean(axis=1))\n    test_data[i] = (test_data[i] - test_pre_train[i, p1:p2].mean(axis=0)) / test_pre_train[i, list(range(p1-30)) + list(range(p2+30, 187))].mean(axis=0) * 1000.0\n\n# Prepare for inference\ntest_x = torch.from_numpy(test_data).unsqueeze(1).float()  # (Number of samples, 1, 375, 283)\ntest_dataset = TensorDataset(test_x)\ntest_loader = DataLoader(test_dataset, batch_size=16, shuffle=False)\n\n#----------------------------------------------------------------------------------------------\n\n# Load the model (using pickle to load the model)\nwith open('/kaggle/input/mk-pickle1/k_fold_4.pickle', 'rb') as f:\n    model = pickle.load(f)\nmodel = model.cuda()  # Move model to GPU\nmodel.eval()\n\n\n# Lists for predictions\npreds = np.zeros((len(test_dataset), 283))\nsigmas = np.zeros((len(test_dataset), 283))\n\n# Make predictions on the test data\nwith torch.no_grad():\n    for i, data in enumerate(test_loader):\n        inputs = data[0].cuda()\n        outputs = model(inputs).reshape((inputs.shape[0], 283, 3))\n        \n        # Store the median (0.5 percentile) in preds and calculate sigma as the upper limit - lower limit\n        preds[i * test_loader.batch_size: (i + 1) * test_loader.batch_size] = outputs[:, :, 1].detach().cpu().numpy()\n        sigmas[i * test_loader.batch_size: (i + 1) * test_loader.batch_size] = (\n            outputs[:, :, 2].detach().cpu().numpy() - outputs[:, :, 0].detach().cpu().numpy()\n        ) \n\n# Prepare the submission file\nss = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/sample_submission.csv')\npreds = preds.clip(0)  # Clip predictions to 0 or above\n\n# Clip sigmas to 0 or above\nsigmas = sigmas.clip(0)\n\n# Concatenate predictions and uncertainties to match the submission format\nsubmission = pd.DataFrame(np.concatenate([preds, sigmas], axis=1), columns=ss.columns[1:])\nsubmission.index = adc_info.index\n\n# Write the submission file\nsubmission.to_csv('submission.csv')\n```",
      "votes": 4
    },
    {
      "id": 2991507,
      "postDate": "2024-09-17T14:19:50.393Z",
      "content": "<p>Not sure if this would help, but try to suppress scientific notation?</p>",
      "rawMarkdown": "Not sure if this would help, but try to suppress scientific notation?",
      "replies": [
        {
          "id": 2991968,
          "postDate": "2024-09-18T03:07:13.997Z",
          "content": "<p>Thank you for your response. Just to confirm, your suggestion was to modify the code like this, correct?</p>\n<p>I tried it, but the same error occurred (perhaps the fix was insufficient).</p>\n<p><code>submission.to_csv('submission.csv', float_format='%.10f')</code></p>",
          "rawMarkdown": "Thank you for your response. Just to confirm, your suggestion was to modify the code like this, correct?\n\nI tried it, but the same error occurred (perhaps the fix was insufficient).\n\n`submission.to_csv('submission.csv', float_format='%.10f')`"
        }
      ]
    },
    {
      "id": 2991319,
      "postDate": "2024-09-17T10:38:52.797Z",
      "content": "<p>I also kept encountering this issue until I avoided pred &lt; 1e-4. Maybe you can give it a try.</p>",
      "rawMarkdown": "I also kept encountering this issue until I avoided pred < 1e-4. Maybe you can give it a try.",
      "replies": [
        {
          "id": 2991971,
          "postDate": "2024-09-18T03:08:41.820Z",
          "content": "<p>Thank you for your response.<br>\nI clipped σ at 0.00011, but the same error occurred…</p>",
          "rawMarkdown": "Thank you for your response.\nI clipped σ at 0.00011, but the same error occurred...",
          "replies": [
            {
              "id": 2992489,
              "postDate": "2024-09-18T15:57:03.113Z",
              "content": "<p>Maybe you can try to clip your pred at 1e-4, not sigma.</p>",
              "rawMarkdown": "Maybe you can try to clip your pred at 1e-4, not sigma.",
              "votes": 1
            },
            {
              "id": 2992532,
              "postDate": "2024-09-18T16:59:36.960Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2992533,
              "postDate": "2024-09-18T16:59:58.263Z",
              "content": "<p>I resolved the issue through other code fixes. It seems the error occurred only with the test data due to the lack of robustness in the phase settings.<br>\nThank you for your help and for considering various solutions.</p>",
              "rawMarkdown": "I resolved the issue through other code fixes. It seems the error occurred only with the test data due to the lack of robustness in the phase settings.\nThank you for your help and for considering various solutions."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2991507,
      "author_name": "ChingYinNg",
      "author_url": "",
      "post_date": "2024-09-17T14:19:50.393000",
      "content": "<p>Not sure if this would help, but try to suppress scientific notation?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2991968,
          "author_name": "Makoto Kine",
          "author_url": "",
          "post_date": "2024-09-18T03:07:13.997000",
          "content": "<p>Thank you for your response. Just to confirm, your suggestion was to modify the code like this, correct?</p>\n<p>I tried it, but the same error occurred (perhaps the fix was insufficient).</p>\n<p><code>submission.to_csv('submission.csv', float_format='%.10f')</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2991319,
      "author_name": "Coolshan",
      "author_url": "",
      "post_date": "2024-09-17T10:38:52.797000",
      "content": "<p>I also kept encountering this issue until I avoided pred &lt; 1e-4. Maybe you can give it a try.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2991971,
          "author_name": "Makoto Kine",
          "author_url": "",
          "post_date": "2024-09-18T03:08:41.820000",
          "content": "<p>Thank you for your response.<br>\nI clipped σ at 0.00011, but the same error occurred…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2992489,
              "author_name": "Coolshan",
              "author_url": "",
              "post_date": "2024-09-18T15:57:03.113000",
              "content": "<p>Maybe you can try to clip your pred at 1e-4, not sigma.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2992532,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-09-18T16:59:36.960000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2992533,
              "author_name": "Makoto Kine",
              "author_url": "",
              "post_date": "2024-09-18T16:59:58.263000",
              "content": "<p>I resolved the issue through other code fixes. It seems the error occurred only with the test data due to the lack of robustness in the phase settings.<br>\nThank you for your help and for considering various solutions.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2990464": "The submission scoring error continues.\nThe code works fine on the provided single test data, and I have confirmed that there are no values below zero (both the predicted values and sigma are clipped at 0, so they won't go below zero).\nI am using the following code for inference.\n\nPlease let me know if you notice anything.\n\n```python\nfrom torch.utils.data import TensorDataset, DataLoader  # Added\n\n# Loading the test data\ndataset = 'test'\nadc_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/'+f'{dataset}_adc_info.csv', index_col='planet_id')\ntest_pre_train = np.concatenate([\n    preproc(f'{dataset}', adc_info, \"FGS1\", 30*12), \n    preproc(f'{dataset}', adc_info, \"AIRS-CH0\", 30)], \n    axis=2\n)\n\n# Generate features using the phase_detector function\ntest_data = test_pre_train.copy()\nfor i in range(len(test_pre_train)):\n    p1, p2 = phase_detector(test_pre_train[i,:,1:].mean(axis=1))\n    test_data[i] = (test_data[i] - test_pre_train[i, p1:p2].mean(axis=0)) / test_pre_train[i, list(range(p1-30)) + list(range(p2+30, 187))].mean(axis=0) * 1000.0\n\n# Prepare for inference\ntest_x = torch.from_numpy(test_data).unsqueeze(1).float()  # (Number of samples, 1, 375, 283)\ntest_dataset = TensorDataset(test_x)\ntest_loader = DataLoader(test_dataset, batch_size=16, shuffle=False)\n\n#----------------------------------------------------------------------------------------------\n\n# Load the model (using pickle to load the model)\nwith open('/kaggle/input/mk-pickle1/k_fold_4.pickle', 'rb') as f:\n    model = pickle.load(f)\nmodel = model.cuda()  # Move model to GPU\nmodel.eval()\n\n\n# Lists for predictions\npreds = np.zeros((len(test_dataset), 283))\nsigmas = np.zeros((len(test_dataset), 283))\n\n# Make predictions on the test data\nwith torch.no_grad():\n    for i, data in enumerate(test_loader):\n        inputs = data[0].cuda()\n        outputs = model(inputs).reshape((inputs.shape[0], 283, 3))\n        \n        # Store the median (0.5 percentile) in preds and calculate sigma as the upper limit - lower limit\n        preds[i * test_loader.batch_size: (i + 1) * test_loader.batch_size] = outputs[:, :, 1].detach().cpu().numpy()\n        sigmas[i * test_loader.batch_size: (i + 1) * test_loader.batch_size] = (\n            outputs[:, :, 2].detach().cpu().numpy() - outputs[:, :, 0].detach().cpu().numpy()\n        ) \n\n# Prepare the submission file\nss = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/sample_submission.csv')\npreds = preds.clip(0)  # Clip predictions to 0 or above\n\n# Clip sigmas to 0 or above\nsigmas = sigmas.clip(0)\n\n# Concatenate predictions and uncertainties to match the submission format\nsubmission = pd.DataFrame(np.concatenate([preds, sigmas], axis=1), columns=ss.columns[1:])\nsubmission.index = adc_info.index\n\n# Write the submission file\nsubmission.to_csv('submission.csv')\n```",
    "2991507": "Not sure if this would help, but try to suppress scientific notation?",
    "2991319": "I also kept encountering this issue until I avoided pred < 1e-4. Maybe you can give it a try."
  }
}