{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":70367,"databundleVersionId":9188054,"sourceType":"competition"},{"sourceId":9806277,"sourceType":"datasetVersion","datasetId":5998787},{"sourceId":204949753,"sourceType":"kernelVersion"}],"dockerImageVersionId":30762,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This code is released under the CC BY 4.0 license, which allows you to use and alter this code (including commercially). You must, however, ensure to give appropriate credit to the original author (Jeroen Cottaar). For details, see https://creativecommons.org/licenses/by/4.0/\n\nThis notebook corresponds to my best submission for the Ariel Data Challenge 2024. It can run in training mode or inference mode. Training mode will typically be run offline, while inference mode will run on Kaggle.\n\nIf you want to get a more visual feel for what the model is up to, head here: https://www.kaggle.com/code/jeroencottaar/ariel-2nd-place-visualization-notebook\n\nFor a general explanation of what is going on here: https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543853\n\nDespite my best efforts, this notebook doesn't seem to use the GPU. Nonetheless it runs faster using the \"GPU T4 x2\" accelerator.\n\nFirst, we set up our toolboxes and Kaggle environment.","metadata":{}},{"cell_type":"code","source":"# Load libraries\nimport sys\nsys.path.append( \"/kaggle/input/my-ariel-library/\" )\nimport ariel_support as ars # general support functions and the data loader\nimport ariel_gp as arg # The model\nimport pandas as pd\nimport numpy as np\nimport copy\nimport dill\nimport IPython","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:45:57.525172Z","iopub.execute_input":"2025-07-25T06:45:57.525404Z","iopub.status.idle":"2025-07-25T06:46:00.66024Z","shell.execute_reply.started":"2025-07-25T06:45:57.525379Z","shell.execute_reply":"2025-07-25T06:46:00.659397Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next, we configure our data loader and model, as it will be used during training. Test set specific tweaks come later. The 'include_later_optimizations' option determines whether we get my original best submission (False) or include later learnings (True).","metadata":{}},{"cell_type":"code","source":"include_later_optimizations = True\n\n# Configure loader\nloader = ars.DataLoader()\nloader.loader_options = ars.baseline_loader(include_later_optimization=include_later_optimizations)\n\n# Configure model\nmodel = arg.baseline_model(include_later_optimization=include_later_optimizations)\n\nif include_later_optimizations:\n    trained_model_filename = 'trained_model_optimized.pickle'\nelse:\n    trained_model_filename = 'trained_model.pickle'","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:46:00.663218Z","iopub.execute_input":"2025-07-25T06:46:00.665789Z","iopub.status.idle":"2025-07-25T06:46:00.676007Z","shell.execute_reply.started":"2025-07-25T06:46:00.66574Z","shell.execute_reply":"2025-07-25T06:46:00.675256Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We're now ready to train our model. By default this is only done if we're not running in the Kaggle environment, but you can alter this in the first line below. If we don't train, we will get our trained model from the earlier stored \"trained_model.pickle\". \n\nWith the default model, only two parameters are trained: the multiplication applied to all sigma's (trained_model.fudge_value), and the multiplication applied to the mean of every transit prediction (trained_model.model.bias). With the optimized model, the training actually does nothing...\n\nThe pickle file for the default model is quite large (~400 MB) because it caches some internal results, which speeds up inference if we do it on the training data later. These caches do not end up being used during inference on test data.","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[]}},{"cell_type":"code","source":"do_training = not ars.running_on_kaggle\nuntrained_model = copy.deepcopy(model)\nif do_training:\n    assert len(ars.test_planet_list)==1, \"Cannot train when submitted\"\n    loader.planet_ids_to_load = ars.train_planet_list\n\n    # Load training data and train model on it\n    train_data = loader.load()\n    model.train(train_data)\n    \n    pickle_data = dict()        \n    pickle_data['untrained_model'] = untrained_model\n    pickle_data['trained_model'] = model\n    ars.pickle_save(ars.file_loc()+trained_model_filename, pickle_data)\npickle_data = ars.pickle_load(ars.file_loc()+trained_model_filename)\nassert dill.dumps(untrained_model) == dill.dumps(pickle_data['untrained_model']), \"trained_model.pickle is not consistent with configured model; perhaps you need to redo training above?\"\ntrained_model = pickle_data['trained_model']","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:46:00.67724Z","iopub.execute_input":"2025-07-25T06:46:00.677879Z","iopub.status.idle":"2025-07-25T06:46:00.734151Z","shell.execute_reply.started":"2025-07-25T06:46:00.677835Z","shell.execute_reply":"2025-07-25T06:46:00.733497Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"If desired we can tweak our model now, to account for differences between the train and test set.","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[]}},{"cell_type":"code","source":"# These values were found by hill climbing the public test set\nif not include_later_optimizations:\n    trained_model.fudge_value += 0.058\n    trained_model.model.bias += -0.0015","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:46:00.735356Z","iopub.execute_input":"2025-07-25T06:46:00.735646Z","iopub.status.idle":"2025-07-25T06:46:00.744053Z","shell.execute_reply.started":"2025-07-25T06:46:00.735615Z","shell.execute_reply":"2025-07-25T06:46:00.741223Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have our trained_model - trained just now or imported - we're ready to apply it to the test data. If we've not actually been submitted, we instead use the first 5 training planets. We can't use the single test planet because some model variations don't work on a single planet.","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[]}},{"cell_type":"code","source":"%%time\n# Load data\nif len(ars.test_planet_list)==1:\n    loader.load_train = True\n    loader.planet_ids_to_load = ars.train_planet_list[:5]\nelse:\n    loader.load_train = False\n    loader.planet_ids_to_load = ars.test_planet_list    \nloader.include_labels = False # don't try to load ground truth\ntest_data = loader.load()\n\n# Do inference - the unused third output has the covariance matrices per planet\npred,sigma,_ = trained_model.infer(test_data)","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:46:00.744767Z","iopub.execute_input":"2025-07-25T06:46:00.745044Z","iopub.status.idle":"2025-07-25T06:48:41.4057Z","shell.execute_reply.started":"2025-07-25T06:46:00.745013Z","shell.execute_reply":"2025-07-25T06:48:41.404265Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Finally, we write the output to submission.csv. If we're not in submission mode we write it to the screen.","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[]}},{"cell_type":"code","source":"# Convert to correct output format\nsubmission = pd.read_csv(ars.data_dir() + '/sample_submission.csv')\nsubmission = submission[0:0]\nfor i in range(len(loader.planet_ids_to_load)):\n    submission.loc[i] = np.concatenate(([loader.planet_ids_to_load[i]], pred[i], sigma[i]))\nsubmission_csv=submission.copy().set_index(\"planet_id\")\nsubmission_csv[submission_csv<=0] = 1e-9\n\n# Output\nif len(ars.test_planet_list)>1:\n    submission_csv.to_csv('submission.csv')\nelse:    \n    IPython.display.display(submission_csv)\n    print(submission_csv.to_numpy()[0,5], submission_csv.to_numpy()[0,500], ars.rms(sigma))\n    # Expected for testing (with include_later_optimizations=False): \n    # Offline: 0.0011164775705400421 3.989177792434888e-05 3.563359762792441e-05\n    # Kaggle:  0.001116477161465638 4.504426029385818e-05 3.567612829731749e-05\n    # Expected for testing (with include_later_optimizations=True): \n    # Offline: 0.0011126207347777167 3.664632827623065e-05 3.197817532055578e-05\n    # Kaggle:  0.0011126215985547824 4.1359960071418506e-05 3.2244350526727034e-05","metadata":{"editable":true,"slideshow":{"slide_type":""},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-07-25T06:48:41.40783Z","iopub.execute_input":"2025-07-25T06:48:41.408351Z","iopub.status.idle":"2025-07-25T06:48:41.619395Z","shell.execute_reply.started":"2025-07-25T06:48:41.408292Z","shell.execute_reply":"2025-07-25T06:48:41.618599Z"}},"outputs":[],"execution_count":null}]}