{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":101849,"databundleVersionId":13093295,"sourceType":"competition"},{"sourceId":13194989,"sourceType":"datasetVersion","datasetId":8353810},{"sourceId":253107285,"sourceType":"kernelVersion"},{"sourceId":255090856,"sourceType":"kernelVersion"}],"dockerImageVersionId":31090,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"See the accompanying solution write-up here: https://www.kaggle.com/competitions/ariel-data-challenge-2025/writeups/1st-place-solution-bayesian-inference-of-course\n\nThis code does not reproduce my winning submission exactly; I cleaned up the code in several ways that slightly improve or degrade the performance. If you want to audit the full original submission, check it out here: https://www.kaggle.com/code/jeroencottaar/ariel2025-1st-place-reproduce-winning-submission\n\nThis final submission code may make it look straightforward. But if you want an idea of what actually goes into this, check the full development history: https://github.com/jcottaar/ariel2\n\nAll implementation is in the attached 'my-ariel2025-library'. Let's start by loading the modules.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"# Load modules\nimport sys\nsys.path.append('/kaggle/input/my-ariel2025-library')\nimport kaggle_support as kgs\nimport ariel_model\nimport ariel_gp\nimport copy","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next, we load all train and test data. This just loads general stuff like planet IDs and initial transit parameters; we don't touch the actual sensor readings here yet.\n\nIf ```fast_mode``` is true, we only use 20 planets from the train set when not submitting to speed things up.","metadata":{}},{"cell_type":"code","source":"train_data = kgs.load_all_train_data()\ntest_data = kgs.load_all_test_data()\nfast_mode = False\nif fast_mode and not kgs.is_submission:\n    train_data = train_data[:20]\nif not kgs.is_submission:\n    # Use train data as test data if not submitting\n    test_data = copy.deepcopy(train_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T20:05:57.883237Z","iopub.execute_input":"2025-09-27T20:05:57.883451Z","iopub.status.idle":"2025-09-27T20:06:04.085426Z","shell.execute_reply.started":"2025-09-27T20:05:57.883433Z","shell.execute_reply":"2025-09-27T20:06:04.084571Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Before running the full model, let's visualize how its various elements behave for one planet.","metadata":{}},{"cell_type":"code","source":"model_visualization = ariel_gp.PredictionModel() # the core Bayesian model; this is not the model we will use for submission    \nmodel_visualization.plot_final = True # make diagnostic plots\nmodel_visualization.train(train_data) # does nothing for this particular model, but mandatory\nmodel_visualization.infer(train_data[0:1]);","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T20:08:32.765136Z","iopub.execute_input":"2025-09-27T20:08:32.765453Z","iopub.status.idle":"2025-09-27T20:10:26.093193Z","shell.execute_reply.started":"2025-09-27T20:08:32.765431Z","shell.execute_reply":"2025-09-27T20:10:26.092313Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now we define and train the full model. The full model adds some stuff to the core Bayesian model above:\n\n- Train the PCA components of the transit depth training labels.\n- Fit several 'fudge factors' to adapt the predicted transit depths and uncertainty margins.\n- Average over multiple transits for one planet.\n\nTraining takes about 4 hours on the full dataset. If you find you lack time in submission, you can pretrain the model offline (saving ```model``` after the ```train``` step).","metadata":{}},{"cell_type":"code","source":"model = ariel_model.baseline_model()\n\n# If you want to try any changes to the model, make them here!\n\nmodel.train(train_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T20:12:09.894829Z","iopub.execute_input":"2025-09-27T20:12:09.895683Z","iopub.status.idle":"2025-09-27T20:16:18.052819Z","shell.execute_reply.started":"2025-09-27T20:12:09.895655Z","shell.execute_reply":"2025-09-27T20:16:18.052079Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Finally, infer on the test data and write the output CSV. If we're not submitting, we also show the local score; note that if ```fast_mode``` above is true this is not representative at all due to the small data size.\n\nInferring takes about 4 hours on the full test set (including the time to load all the sensor data).","metadata":{}},{"cell_type":"code","source":"inferred_data = model.infer(test_data)\nif not kgs.is_submission:\n    print(kgs.score_metric(inferred_data, test_data))\ndf = kgs.make_submission_dataframe(inferred_data)\nkgs.write_submission_csv(df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T20:16:18.054038Z","iopub.execute_input":"2025-09-27T20:16:18.054279Z","iopub.status.idle":"2025-09-27T20:19:33.605602Z","shell.execute_reply.started":"2025-09-27T20:16:18.05426Z","shell.execute_reply":"2025-09-27T20:19:33.60471Z"}},"outputs":[],"execution_count":null}]}