{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70367,"databundleVersionId":9188054,"sourceType":"competition"},{"sourceId":217766477,"sourceType":"kernelVersion"},{"sourceId":217834937,"sourceType":"kernelVersion"}],"dockerImageVersionId":30839,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Simplified and Concise Explanation of the Challenge\n## Objective:\nTo extract the chemical spectrum of exoplanet atmospheres using simulated data from the ARIEL mission, which includes observations of planetary transits.\nARIEL Mission: Observes exoplanets as they transit their stars.\n### Dataset Overview:\nData collected includes time-series imagery from two instruments:\n  * FGS1 (Fine Guidance System 1): Visible light (0.60–0.80 µm).\n  * AIRS-CH0 (Infrared Spectrometer): Infrared (1.95–3.90 µm).\n    \nKey Characteristics:\n* Data includes noise and is based on limited photon counts.\n* Each planetary observation produces a large volume of images:\n* FGS1: 135,000 frames, each 32x32 pixels.\n* AIRS-CH0: 11,250 frames, each 32x356 pixels.\n* Data files store flattened images requiring reshaping for analysis.\n  \nHidden Test Set:\n* Approximately 800 exoplanets in the test set.\n* Includes some real exoplanet simulations (not scored).\n  \n### Data Components:\nSignal Files:\n* Time-series flux measurements during transits.\n* Stored as .parquet files for both instruments (FGS1 and AIRS-CH0).\n\nCalibration Files: Used to correct raw sensor data for noise and sensor irregularities.\\\nTypes of calibration files:\n* Dark Frames: Subtract thermal noise and bias levels.\n* Dead Frames: Identify unresponsive (dead) or overly responsive (hot) pixels.\n* Flat Frames: Correct pixel sensitivity variations.\n* Linearity Correction: Adjust for non-linear pixel response at high charges.\n* Read Noise Frames: Account for electronic noise during readout.\n       \nMetadata Files:\n* ADC Info: Restores the original dynamic range of the data using gain and offset values.\n* Axis Info: Provides temporal and spectral axis details for the time-series data.\n* Wavelength Grid: Links spectra to wavelength ranges.\n* Train Labels: Ground truth chemical spectra for the training data.\n","metadata":{}},{"cell_type":"markdown","source":"# First Step : Calibration \nThe code for calibration is in the notebook: https://www.kaggle.com/code/irasharma/neurips/notebook?scriptVersionId=217626513 \nwhich includes the following steps:\n\n1. ADC Conversion: Convert digital signals back to analog form using gain and offset.\n2. Bad Pixel Correction: Mask and correct hot and dead pixels using provided maps and dark frame data.\n3. Non-Linearity Correction: Apply non-linearity corrections using spline interpolation for better signal accuracy.\n4. Dark Current Removal: Subtract dark current while accounting for dead pixels.\n5. CDS Calculation: Perform Correlated Double Sampling (CDS) by subtracting the start-of-exposure frames from the end-of-exposure frames.\n6. Time Binning: Bin the data over specified time intervals for more efficient analysis.\n7. Flat Field Correction: Correct signal using flat field data to adjust for pixel-to-pixel variations in the detector.\n8. Index Extraction: Retrieve the indices of training data using file paths and split into chunks.\n9. Chunk Processing: Process the data in chunks, applying various corrections (masking, non-linearity, dark current, flat field) based on configuration settings.\n10. Data Preprocessing: Read and preprocess FGS1 and AIRS-CH0 signal and calibration data, applying corrections as specified (masking, dark current removal, etc.).\n11. Final Output: Cleaned and calibrated signals as AIRS_clean_train_{planet_id}.npy and FGS_clean_train_{planet_id}.npy are stored for all 673 planets for further analysis or inference.\n","metadata":{}},{"cell_type":"markdown","source":"# Second Step: Analysis of the Calibrated FGS1 and AIRS-CHO Signals","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport polars as pl\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport seaborn as sns\nimport scipy.stats\nfrom tqdm import tqdm\nimport pickle\nimport itertools\nimport os\nimport glob\nfrom astropy.stats import sigma_clip\nfrom scipy.optimize import curve_fit\nfrom sklearn.model_selection import train_test_split, cross_val_predict\nfrom sklearn.linear_model import Ridge, RidgeCV, LinearRegression\nfrom sklearn.metrics import r2_score, mean_squared_error\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:09:18.286313Z","iopub.execute_input":"2025-01-16T10:09:18.286557Z","iopub.status.idle":"2025-01-16T10:09:22.568234Z","shell.execute_reply.started":"2025-01-16T10:09:18.286533Z","shell.execute_reply":"2025-01-16T10:09:22.566901Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_labels=pd.read_csv(\"/kaggle/input/ariel-data-challenge-2024/train_labels.csv\")\ntrain_labels","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:09:26.19331Z","iopub.execute_input":"2025-01-16T10:09:26.193665Z","iopub.status.idle":"2025-01-16T10:09:26.341983Z","shell.execute_reply.started":"2025-01-16T10:09:26.193639Z","shell.execute_reply":"2025-01-16T10:09:26.340979Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\n\n# Define the path to the dataset directory\ndataset_path = '/kaggle/input/long-run-neurips/data_light_raw'\n\n# List all files in the directory\nfiles = os.listdir(dataset_path)\n\n# Filter out the files that end with 'AIRS_clean_train_' and '.npy'\nairs_clean_files = [file for file in files if file.startswith('AIRS_clean_train_')]\nfgs_clean_files = [file for file in files if file.startswith('FGS1_train_')]\n# Display the number of matching files\nprint(f\"Number of AIRS_clean_train_ files: {len(airs_clean_files)}\")\nprint(f\"Number of FGS_clean_train_ files: {len(fgs_clean_files)}\")\n# print(f\"Files found: {airs_clean_files}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:09:30.047095Z","iopub.execute_input":"2025-01-16T10:09:30.047483Z","iopub.status.idle":"2025-01-16T10:09:30.101304Z","shell.execute_reply.started":"2025-01-16T10:09:30.047452Z","shell.execute_reply":"2025-01-16T10:09:30.099724Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### AIRS-CH0 Calibrated Signal Visualization","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n\nplt.figure(figsize=(10, 6))\n\n# Iterate over the first 5 planets\nfor planet_id in train_labels['planet_id'].values[:1]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/AIRS_clean_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Plot the heatmap on top of others\n    sns.heatmap(data[0, :, :].T, cbar=True)\n\n# Customize the plot\nplt.title(\"Intensity Heatmap for a Planet\")\nplt.ylabel('Spatial Dimension')\nplt.xlabel('Wavelength Dimension')\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:09:33.500417Z","iopub.execute_input":"2025-01-16T10:09:33.50075Z","iopub.status.idle":"2025-01-16T10:09:34.6Z","shell.execute_reply.started":"2025-01-16T10:09:33.500723Z","shell.execute_reply":"2025-01-16T10:09:34.598833Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Brightness Variation Information Showing the Approximate In-Transit and Out-Transit Regions\n\nplt.figure(figsize=(10, 6))\n\n# Iterate for all the planets\nfor planet_id in train_labels['planet_id'].values[:]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/AIRS_clean_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Sum over the wavelength \n    series = data.sum(axis=1).sum(axis=1)\n\n    plt.plot(series/series.mean())\n\nfor time_step in [51, 75, 115, 140]:\n   plt.axvline(time_step, color='black', linewidth=1)\nplt.title(\"For 673 Planets\")\nplt.xlabel(\"Time (frame index)\")\nplt.ylabel(\"Normalized Flux\")\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:09:37.443532Z","iopub.execute_input":"2025-01-16T10:09:37.443906Z","iopub.status.idle":"2025-01-16T10:10:49.680697Z","shell.execute_reply.started":"2025-01-16T10:09:37.443839Z","shell.execute_reply":"2025-01-16T10:10:49.679479Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Smoothened Gradient showing the Approximate Transition Points\n\nplt.figure(figsize=(10, 6))\n\n# Iterate for all the planets\nfor planet_id in train_labels['planet_id'].values[:]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/AIRS_clean_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Sum over the wavelength \n    tsignal=pd.Series(data.sum(axis=1).sum(axis=1)).rolling(window=5).sum()\n    gsignal=np.gradient(tsignal/tsignal.max())\n\n    plt.plot(gsignal)\n\nfor time_step in [51, 75, 115, 140]:\n   plt.axvline(time_step, color='black', linewidth=1)\nplt.title(\"For 673 Planets\")\nplt.xlabel(\"Time (frame index)\")\nplt.ylabel(\"Normalized Gradient Flux\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:10:49.682209Z","iopub.execute_input":"2025-01-16T10:10:49.682496Z","iopub.status.idle":"2025-01-16T10:10:58.514966Z","shell.execute_reply.started":"2025-01-16T10:10:49.682472Z","shell.execute_reply":"2025-01-16T10:10:58.513596Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Spectral Information showing the absorption dips\n\nplt.figure(figsize=(10, 6))\n\n# Iterate for all the planets\nfor planet_id in train_labels['planet_id'].values[:]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/AIRS_clean_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Sum over the time frame\n    series = data.sum(axis=0).sum(axis=1)\n\n    plt.plot(series/series.mean())\n\n# Customize the plot\nplt.title(\"For 673 Planets\")\nplt.xlabel(\"Wavelength\")\nplt.ylabel(\"Normalized Flux\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:10:58.517245Z","iopub.execute_input":"2025-01-16T10:10:58.517546Z","iopub.status.idle":"2025-01-16T10:11:05.439246Z","shell.execute_reply.started":"2025-01-16T10:10:58.517522Z","shell.execute_reply":"2025-01-16T10:11:05.438038Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### FGS1 Calibrated Signal Visualization","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n\nplt.figure(figsize=(10, 6))\n\n# Iterate over the first 5 planets\nfor planet_id in train_labels['planet_id'].values[:1]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/FGS1_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Plot the heatmap on top of others\n    sns.heatmap(data[0, :, :].T, cbar=True)\n\n# Customize the plot\nplt.title(\"Intensity Heatmap for a Planet\")\nplt.ylabel('Spatial Dimension')\nplt.xlabel('Wavelength Dimension')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:11:05.440812Z","iopub.execute_input":"2025-01-16T10:11:05.441194Z","iopub.status.idle":"2025-01-16T10:11:06.212077Z","shell.execute_reply.started":"2025-01-16T10:11:05.441166Z","shell.execute_reply":"2025-01-16T10:11:06.210978Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Brightness Variation Information Showing the Approximate In-Transit and Out-Transit Regions\n\nplt.figure(figsize=(10, 6))\n\n# Iterate for all the planets\nfor planet_id in train_labels['planet_id'].values[:]:\n    # Construct the file path\n    file_path = f\"/kaggle/input/long-run-neurips/data_light_raw/FGS1_train_{planet_id}.npy\"\n    \n    # Load the data for the planet\n    data = np.load(file_path)[0]\n    \n    # Replace NaN values with 0\n    data[np.isnan(data)] = 0\n    \n    # Sum over the wavelength \n    series = data.sum(axis=1).sum(axis=1)\n\n    plt.plot(series/series.mean())\n\nfor time_step in [51, 75, 115, 140]:\n   plt.axvline(time_step, color='black', linewidth=1)\nplt.title(\"For 673 Planets\")\nplt.xlabel(\"Time (frame index)\")\nplt.ylabel(\"Normalized Flux\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:11:06.213225Z","iopub.execute_input":"2025-01-16T10:11:06.213614Z","iopub.status.idle":"2025-01-16T10:11:22.746287Z","shell.execute_reply.started":"2025-01-16T10:11:06.213578Z","shell.execute_reply":"2025-01-16T10:11:22.745153Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Third Step : Transit Analysis\n\nOnly the AIRS-CH0 data is used for the Transit analysis because it encompasses both the required brightness intensity dips and spectral information. While the FGS1 data, which solely provides brightness variation details, is just supposed to be added as a single column later.\n\nThe code to find Transit Depth for all the 673 planets is in the notebook: https://www.kaggle.com/code/irasharma/transitdata-errors/notebook?scriptVersionId=217834937 which includes the following steps:\n\n File Management:\n   * Extracts dataset IDs from files containing 'AIRS_clean_train_'.\n   * Loads data files and replaces NaN values with 0.\n     \n Light Curve Extraction:\n   * Computes the total light intensity by summing data across axes.\n   * Normalizes the light curve by its mean intensity.\n     \n Detrending Using Taylor Series:\n\n   * Models light curve trends using a Taylor series function: $ f(x) = a + b \\cdot x + c \\cdot x^2 + d \\cdot x^3 $\n   * Fits the function to the start and end regions of the light curve using curve_fit.\n   * Detrends the light curve by dividing by the fitted model.\n\nTransit Phase Identification:\n   * Computes gradient to locate transit boundaries\n   * Ingress (x_start): Minimum gradient.\n   * Egress (x_end): Maximum gradient.\n\n Spectral Analysis:\n   * Out-of-Transit Phases: Before ingress and after egress.\n   * In-Transit Phase: Between ingress and egress.\n   * Extracts and sums spectra for each phase.\n\n Output Management:\n   * In-transit: /AIRS_spectrum_in_transit/.\n   * Out-of-transit: /AIRS_spectrum_out_transit/.","metadata":{}},{"cell_type":"markdown","source":"# Fourth Step: Regression Training Model","metadata":{}},{"cell_type":"code","source":"in_data=np.load(\"/kaggle/input/transitdata-errors/1005054328/in_transit_spectrum_data.npy\")\nout_data= np.load(\"/kaggle/input/transitdata-errors/1005054328/out_transit_spectrum_data.npy\")\nplanet_id = 1005054328  # Example planet ID as an integer\n\n# Filter the DataFrame to get the row corresponding to this planet_id\nplanet_data = train_labels[train_labels['planet_id'] == planet_id]\n\n# Drop the 'planet_id' column, as you only need the wavelength columns\nwavelength_data = planet_data.drop(columns=['planet_id'])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:11:22.74733Z","iopub.execute_input":"2025-01-16T10:11:22.747604Z","iopub.status.idle":"2025-01-16T10:11:22.772781Z","shell.execute_reply.started":"2025-01-16T10:11:22.747581Z","shell.execute_reply":"2025-01-16T10:11:22.771691Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tr_data = []\n# Iterate over the first 5 planets\nfor planet_id in tqdm(train_labels['planet_id'].values[:]):\n    # Construct the file path\n    in_data = np.load(f\"/kaggle/input/transitdata-errors/{planet_id}/in_transit_spectrum_data.npy\")\n    out_data = np.load(f\"/kaggle/input/transitdata-errors/{planet_id}/out_transit_spectrum_data.npy\")\n    transit_depth= (out_data-in_data)/out_data\n    tr_data.append(transit_depth)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:11:22.773997Z","iopub.execute_input":"2025-01-16T10:11:22.77437Z","iopub.status.idle":"2025-01-16T10:11:34.804678Z","shell.execute_reply.started":"2025-01-16T10:11:22.774342Z","shell.execute_reply":"2025-01-16T10:11:34.803619Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"tr_data = np.array(tr_data)\ntrain_data, val_data, train_label, val_label = train_test_split(tr_data, train_labels.values[:, 1:], \\\n                                                               test_size=0.4, shuffle=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:14:41.077514Z","iopub.execute_input":"2025-01-16T10:14:41.079033Z","iopub.status.idle":"2025-01-16T10:14:41.09745Z","shell.execute_reply.started":"2025-01-16T10:14:41.078973Z","shell.execute_reply":"2025-01-16T10:14:41.09615Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.linear_model import Ridge\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.metrics import r2_score, mean_squared_error\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# Define Ridge regression model\nmodel = Ridge(alpha=1E-3)\n\n# Cross-validation predictions for train data\noof_pred = cross_val_predict(model, train_data, train_label)\n\n# Calculate R2 score and RMSE for train data\nprint(f\"# R2 score for Train Data: {r2_score(train_label, oof_pred):.3f}\")\nsigma_pred = mean_squared_error(train_label, oof_pred, squared=False)\nprint(f\"# Root mean squared error for Train Data: {sigma_pred:.6f}\")\n\n# Cross-validation predictions for validation data\noof_pred_val = cross_val_predict(model, val_data, val_label)\n\n# Calculate R2 score and RMSE for validation data\nprint(f\"# R2 score on Validation Data: {r2_score(val_label, oof_pred_val):.3f}\")\nsigma_pred_val = mean_squared_error(val_label, oof_pred_val, squared=False)\nprint(f\"# Root mean squared error for Validation Data: {sigma_pred_val:.6f}\")\n\n# Plotting\nplt.figure(figsize=(8, 6))\ncol = 1  # Ensure this column exists in both train_label and val_label\n\n# Error bar plot for train data\nplt.errorbar(\n    oof_pred[:, col], train_label[:, col], yerr=sigma_pred, \n    fmt='o', markersize=4, alpha=1, label='Predictions on Train Data', \n    color='deepskyblue', markeredgecolor='blue', ecolor='gray', elinewidth=2\n)\n\n# Error bar plot for validation data\nplt.errorbar(\n    oof_pred_val[:, col], val_label[:, col], yerr=sigma_pred_val, \n    fmt='o', markersize=4, alpha=1, label='Predictions on Validation Data', \n    color='lightgreen', markeredgecolor='green', ecolor='gray', elinewidth=2\n)\n\n# Reference line y = x\ny_min = min(train_label.min(), val_label.min())\ny_max = max(train_label.max(), val_label.max())\nplt.plot([y_min, y_max], [y_min, y_max], color='red', linestyle='--', linewidth=3, label='y = x')\n\n# Labels, title, and legend\nplt.xlabel('y_pred')\nplt.ylabel('y_true')\nplt.title('Comparing y_true and y_pred')\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-01-16T10:27:32.883231Z","iopub.execute_input":"2025-01-16T10:27:32.883622Z","iopub.status.idle":"2025-01-16T10:27:33.289057Z","shell.execute_reply.started":"2025-01-16T10:27:32.883591Z","shell.execute_reply":"2025-01-16T10:27:33.287752Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pickle\n\n# save\nwith open('/kaggle/working/model.pkl','wb') as f:\n    pickle.dump(model,f)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}