{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":13093295,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is focused on understanding the structure and content of the Ariel Data Challenge dataset. We begin by loading the main data files, visualizing sample spectra, and inspecting the raw detector images. We then process the data into a standardized format to be used in downstream processing and compare the original data with the processed data. Most of the processing code is from https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data. \n\nChanges: \n- less time binning, only enough for a standard length between the two signals.\n- changes in masking behavior\n- including inpainting\n- visualizations of the signals before and after processing","metadata":{}},{"cell_type":"markdown","source":"***Loading and Visualizing Transmission Spectra***\n\nWe start by loading the transmission spectra for the training set and the corresponding wavelengths. To get a sense of the data, we plot the raw spectra for three example exoplanets, then repeat the visualization after Z-normalizing each spectrum. This helps us understand the typical scale, shape, and variability of the measurements across wavelengths.","metadata":{}},{"cell_type":"code","source":"# preprocessing code borrowed from: https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\nimport numpy as np\nimport pandas as pd\nimport itertools\nimport os\nimport glob \nfrom astropy.stats import sigma_clip\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\nimport os\nimport re\nfrom skimage.restoration import inpaint_biharmonic","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:44:46.87015Z","iopub.execute_input":"2025-09-30T18:44:46.870585Z","iopub.status.idle":"2025-09-30T18:44:47.673324Z","shell.execute_reply.started":"2025-09-30T18:44:46.870554Z","shell.execute_reply":"2025-09-30T18:44:47.672321Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load spectra and wavelengths\ntrain = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train.csv')\nwavelengths = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/wavelengths.csv')\n\n# Columns for spectrum (excluding 'planet_id')\nspectrum_col_names = train.columns.drop('planet_id')\nwavelength_values = wavelengths.loc[0, spectrum_col_names].values.astype(float)  # Shape: (283,)\n\n# Planet indices or planet_ids to compare (first 3 rows by default)\nindices = [0, 1, 2]\n\n# Plot: Original spectra\nplt.figure(figsize=(10, 5))\nfor idx in indices:\n    row = train.iloc[idx]\n    planet_id = row['planet_id']\n    spectrum_values = row[spectrum_col_names].values.astype(float)\n    plt.plot(wavelength_values, spectrum_values, marker='o', label=f'Planet {planet_id}')\nplt.xlabel('Wavelength (μm)')\nplt.ylabel('rp/rs (planet-to-star radius ratio)')\nplt.title('Transmission Spectra for 3 Exoplanets (Raw)')\nplt.legend()\nplt.grid(True)\nplt.tight_layout()\nplt.show()\n\n# Plot: Z-normalized spectra\nplt.figure(figsize=(10, 5))\nfor idx in indices:\n    row = train.iloc[idx]\n    planet_id = row['planet_id']\n    spectrum_values = row[spectrum_col_names].values.astype(float)\n    # Z-normalization: (value - mean) / std\n    z_norm_spectrum = (spectrum_values - spectrum_values.mean()) / spectrum_values.std()\n    plt.plot(wavelength_values, z_norm_spectrum, marker='o', label=f'Planet {planet_id}')\nplt.xlabel('Wavelength (μm)')\nplt.ylabel('Z-normalized rp/rs')\nplt.title('Transmission Spectra for 3 Exoplanets (Z-Normalized)')\nplt.legend()\nplt.grid(True)\nplt.tight_layout()\nplt.show()\n\nprint(spectrum_values.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:44:49.118522Z","iopub.execute_input":"2025-09-30T18:44:49.118948Z","iopub.status.idle":"2025-09-30T18:44:49.93106Z","shell.execute_reply.started":"2025-09-30T18:44:49.118922Z","shell.execute_reply":"2025-09-30T18:44:49.930242Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The first plot shows the raw transmission spectra, which shows that the transmissions have different average values.\n\nThe second plot uses Z-normalization to highlight relative differences and make spectra more comparable across planets.","metadata":{}},{"cell_type":"markdown","source":"***Inspecting Raw Detector Images***\n\nNext, we examine the raw detector images for both AIRS-CH0 and FGS1 instruments. These images are loaded, calibrated using ADC parameters, and visualized for a sample planet. This step helps us understand the data format and calibration process for the pixel-level signals. Calibration is performed using gain and offset values from the ADC info file.","metadata":{}},{"cell_type":"code","source":"# Load ADC calibration parameters (identical for all planets)\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]   # Gain for AIRS-CH0 detector\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0] # Offset for AIRS-CH0 detector\nfgs_gain = adc['FGS1_adc_gain'].iloc[0]         # Gain for FGS1 detector\nfgs_offset = adc['FGS1_adc_offset'].iloc[0]     # Offset for FGS1 detector\n\n# --- AIRS-CH0 IMAGE (32x356), FIRST FRAME ---\n# Load AIRS-CH0 signal data for a sample planet (first time frame)\nairs_path = '/kaggle/input/ariel-data-challenge-2025/train/1010375142/AIRS-CH0_signal_0.parquet'\nairs_signals = pd.read_parquet(airs_path)\nairs_img = airs_signals.iloc[0].values.reshape(32, 356)  # Reshape to detector dimensions\n# Calibrate raw pixel values using gain and offset\nairs_img_cal = (airs_img / airs_gain) + airs_offset\n\n# Visualize calibrated AIRS-CH0 image\nplt.figure(figsize=(12,4))\nplt.imshow(airs_img_cal, cmap='gray', aspect='auto')\nplt.colorbar(label='Calibrated pixel value')\nplt.title('AIRS-CH0 Calibrated Image (First Time Frame)')\nplt.xlabel('Wavelength axis (column index)')\nplt.ylabel('Spatial axis (row index)')\nplt.show()\n\n# --- FGS1 IMAGE (32x32), FIRST FRAME ---\n# Load FGS1 signal data for the same planet (first time frame)\nfgs_path = '/kaggle/input/ariel-data-challenge-2025/train/1010375142/FGS1_signal_0.parquet'\nfgs_signals = pd.read_parquet(fgs_path)\nfgs_img = fgs_signals.iloc[0].values.reshape(32, 32)  # Reshape to detector dimensions\n# Calibrate raw pixel values using gain and offset\nfgs_img_cal = (fgs_img / fgs_gain) + fgs_offset\n\n# Visualize calibrated FGS1 image\nplt.figure(figsize=(6,6))\nplt.imshow(fgs_img_cal, cmap='gray')\nplt.colorbar(label='Calibrated pixel value')\nplt.title('FGS1 Calibrated Image (First Time Frame)')\nplt.xlabel('Pixel X')\nplt.ylabel('Pixel Y')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:44:54.833402Z","iopub.execute_input":"2025-09-30T18:44:54.833722Z","iopub.status.idle":"2025-09-30T18:44:57.72339Z","shell.execute_reply.started":"2025-09-30T18:44:54.833694Z","shell.execute_reply":"2025-09-30T18:44:57.722387Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***Checking Data Shapes and File Structure***\n\nTo ensure consistency and understand the data organization, we check the shapes of AIRS-CH0 and FGS1 signal files for a few sample planets. This step is useful for debugging and for planning downstream processing steps.\n","metadata":{}},{"cell_type":"code","source":"base_dir = '/kaggle/input/ariel-data-challenge-2025/train'\n\n# List of sample planet IDs to check\nsample_planets = ['1010375142', '1024292144', '1029552010']\n\nfor planet_id in sample_planets:\n    # Construct full paths to AIRS-CH0 and FGS1 signal files for each planet\n    airs_path = os.path.join(base_dir, planet_id, 'AIRS-CH0_signal_0.parquet')\n    fgs1_path = os.path.join(base_dir, planet_id, 'FGS1_signal_0.parquet')\n    \n    # Check if AIRS-CH0 file exists, then load and print its shape\n    if os.path.exists(airs_path):\n        airs_data = pd.read_parquet(airs_path)\n        print(f'Planet {planet_id} AIRS-CH0 shape:', airs_data.shape)\n        # Expected shape: (num_time_frames, num_pixels)\n    else:\n        print(f'Planet {planet_id} AIRS-CH0 file missing.')\n    \n    # Check if FGS1 file exists, then load and print its shape\n    if os.path.exists(fgs1_path):\n        fgs1_data = pd.read_parquet(fgs1_path)\n        print(f'Planet {planet_id} FGS1 shape:', fgs1_data.shape)\n        # Expected shape: (num_time_frames, num_pixels)\n    else:\n        print(f'Planet {planet_id} FGS1 file missing.')\n    \n    print('---')  # Separator for readability","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:44:57.784915Z","iopub.execute_input":"2025-09-30T18:44:57.785294Z","iopub.status.idle":"2025-09-30T18:45:02.251098Z","shell.execute_reply.started":"2025-09-30T18:44:57.785263Z","shell.execute_reply":"2025-09-30T18:45:02.249957Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***Visualizing Multiple AIRS-CH0 Frames***\n\nTo further explore the temporal dimension of the AIRS-CH0 data, we plot several consecutive frames for a sample planet. This helps us understand how the signal evolves over time and spot any anomalies or patterns in the raw images.","metadata":{}},{"cell_type":"code","source":"# Load ADC calibration parameters for AIRS-CH0 detector\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]   # Gain for AIRS-CH0\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0] # Offset for AIRS-CH0\n\n# Select a sample planet to visualize\nplanet_id = '1010375142'  # Replace with any valid planet_id as needed\n\n# Load AIRS-CH0 signal data for the selected planet\npath = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_0.parquet'\ndf = pd.read_parquet(path)\n\nframes_to_plot = 9  # Number of time frames to visualize (first 9)\nplt.figure(figsize=(15, 10))\nfor idx in range(frames_to_plot):\n    # Extract and reshape the signal for the current frame (32 rows x 356 columns)\n    img = df.iloc[idx].values.reshape(32, 356)\n    # Calibrate the raw pixel values using gain and offset\n    cal_img = (img / airs_gain) + airs_offset\n    # Plot the calibrated image in a 3x3 grid\n    plt.subplot(3, 3, idx + 1)\n    plt.imshow(cal_img, cmap='gray', aspect='auto')\n    plt.title(f'AIRS Frame {idx}')\n    plt.axis('off')  # Hide axis ticks for clarity\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:45:03.061303Z","iopub.execute_input":"2025-09-30T18:45:03.06164Z","iopub.status.idle":"2025-09-30T18:45:05.059008Z","shell.execute_reply.started":"2025-09-30T18:45:03.06161Z","shell.execute_reply":"2025-09-30T18:45:05.057803Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***Preprocessing and Cleaning Detector Data***\n\nThis section details the preprocessing pipeline for cleaning and calibrating the raw detector signals. \n\nThe pipeline includes masking hot/dead pixels, applying dark and flat field corrections, binning, and inpainting. Most of the code is adapted from this notebook (https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data), with additional comments, structure, and performance improvements for clarity.\n\nThe main changes in this version are changed masked array behavior and the inclusion of inpainting to fill in masked data values to prepare for neural network training.","metadata":{}},{"cell_type":"code","source":"# Path to the folder containing the input data\npath_folder = '/kaggle/input/ariel-data-challenge-2025'\n# Path to the folder where processed light data will be stored\npath_out = '/kaggle/working/data_light_raw/'\noutput_dir = '/kaggle/working/data_light_raw/'  # (redundant, but kept for clarity)\n\n# Check if the output directory exists; create it if it does not\nif not os.path.exists(path_out):\n    os.makedirs(path_out)  # Creates all intermediate directories as needed\n    print(f\"Directory {path_out} created.\")\nelse:\n    print(f\"Directory {path_out} already exists.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:45:06.305564Z","iopub.execute_input":"2025-09-30T18:45:06.305859Z","iopub.status.idle":"2025-09-30T18:45:06.313385Z","shell.execute_reply.started":"2025-09-30T18:45:06.305836Z","shell.execute_reply":"2025-09-30T18:45:06.311928Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"CHUNKS_SIZE = 1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:45:06.940901Z","iopub.execute_input":"2025-09-30T18:45:06.941298Z","iopub.status.idle":"2025-09-30T18:45:06.946148Z","shell.execute_reply.started":"2025-09-30T18:45:06.941269Z","shell.execute_reply":"2025-09-30T18:45:06.945196Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def ADC_convert(signal, gain, offset):\n    \"\"\"\n    Convert raw detector signal to calibrated values using gain and offset.\n    \n    Parameters:\n        signal (ndarray): Raw pixel values.\n        gain (float): Calibration gain from ADC info.\n        offset (float): Calibration offset from ADC info.\n    \n    Returns:\n        ndarray: Calibrated signal as float64.\n    \"\"\"\n    signal = signal.astype(np.float64)\n    signal /= gain\n    signal += offset\n    return signal\n\ndef mask_hot_dead(signal, dead, dark):\n    \"\"\"\n    Mask hot and dead pixels in the signal array.\n    Hot pixels are identified using sigma clipping on the dark frame.\n    Masks are repeated along the time axis to match signal shape.\n    \n    Parameters:\n        signal (ndarray): Detector data (time, x, y).\n        dead (ndarray): Boolean mask of dead pixels (True = dead).\n        dark (ndarray): Dark frame used to identify hot pixels.\n    \n    Returns:\n        np.ma.MaskedArray: Signal with hot and dead pixels masked.\n    \"\"\"\n    hot = sigma_clip(dark, sigma=5, maxiters=5).mask\n    hot = np.tile(hot, (signal.shape[0], 1, 1))\n    dead = np.tile(dead, (signal.shape[0], 1, 1))\n    signal = np.ma.masked_where(dead, signal)\n    signal = np.ma.masked_where(hot, signal)\n    return signal\n\ndef apply_linear_corr(linear_corr, clean_signal):\n    \"\"\"\n    Apply a per-pixel linearity correction to the signal using polynomial coefficients.\n    \n    Parameters:\n        linear_corr (ndarray): Polynomial coefficients (degree, x, y).\n        clean_signal (ndarray): Detector data (time, x, y).\n    \n    Returns:\n        ndarray: Linearity-corrected signal.\n    \"\"\"\n    linear_corr = np.flip(linear_corr, axis=0)\n    for x, y in itertools.product(\n                range(clean_signal.shape[1]), range(clean_signal.shape[2])\n            ):\n        poli = np.poly1d(linear_corr[:, x, y])\n        clean_signal[:, x, y] = poli(clean_signal[:, x, y])\n    return clean_signal\n\ndef clean_dark(signal, dead, dark, dt):\n    \"\"\"\n    Subtract dark current from the signal, accounting for dead pixels and integration time.\n    \n    Parameters:\n        signal (ndarray): Detector data (time, x, y).\n        dead (ndarray): Boolean mask of dead pixels.\n        dark (ndarray): Dark frame (x, y).\n        dt (ndarray): Integration time per frame (time,).\n    \n    Returns:\n        ndarray: Dark-corrected signal.\n    \"\"\"\n    dark = np.ma.masked_where(dead, dark)\n    dark = np.tile(dark, (signal.shape[0], 1, 1))\n    signal -= dark * dt[:, np.newaxis, np.newaxis]\n    return signal\n\ndef get_cds(signal):\n    \"\"\"\n    Compute correlated double sampling (CDS) difference along the time axis.\n    \n    Parameters:\n        signal (ndarray): Detector data (time, ...).\n    \n    Returns:\n        ndarray: CDS-differenced signal (half the time dimension).\n    \"\"\"\n    cds = signal[:,1::2,:,:] - signal[:,::2,:,:]\n    return cds\n\ndef correct_flat_field(flat, dead, signal):\n    \"\"\"\n    Apply flat field correction to the signal.\n    Masks dead pixels in the flat field and repeats along the time axis.\n    \n    Parameters:\n        flat (ndarray): Flat field frame (x, y).\n        dead (ndarray): Boolean mask of dead pixels (x, y).\n        signal (ndarray): Detector data (time, x, y).\n    \n    Returns:\n        ndarray: Flat-field corrected signal.\n    \"\"\"\n    flat = flat.transpose(1, 0)\n    dead = dead.transpose(1, 0)\n    flat = np.ma.masked_where(dead, flat)\n    flat = np.tile(flat, (signal.shape[0], 1, 1))\n    signal = signal / flat\n    return signal","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:45:08.743564Z","iopub.execute_input":"2025-09-30T18:45:08.743863Z","iopub.status.idle":"2025-09-30T18:45:08.756442Z","shell.execute_reply.started":"2025-09-30T18:45:08.743838Z","shell.execute_reply":"2025-09-30T18:45:08.755418Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def get_index(files, CHUNKS_SIZE):\n    \"\"\"\n    Extract planet indices from AIRS-CH0_signal_0.parquet files and split into chunks.\n\n    Parameters:\n        files (list of str): List of file paths.\n        CHUNKS_SIZE (int): Number of indices per chunk.\n\n    Returns:\n        list of np.ndarray: List of arrays, each containing planet indices for a chunk.\n    \"\"\"\n    index = []\n    for file in files:\n        file_name = file.split('/')[-1]\n        if file_name.split('_')[0] == 'AIRS-CH0' and file_name.split('_')[-1] == '0.parquet':\n            file_index = os.path.basename(os.path.dirname(file))\n            index.append(int(file_index))\n    index = np.array(index)\n    index = np.sort(index)\n    # Split indices into chunks of size CHUNKS_SIZE\n    index = np.array_split(index, len(index)//CHUNKS_SIZE)\n    return index\n\n\ndef get_multiobs_index(files, CHUNKS_SIZE):\n    \"\"\"\n    Extract (planet_id, obs_num) pairs from AIRS-CH0_signal_X.parquet files and split into chunks.\n\n    Parameters:\n        files (list of str): List of file paths.\n        CHUNKS_SIZE (int): Number of pairs per chunk.\n\n    Returns:\n        list of list: List of (planet_id, obs_num) tuples, split into chunks.\n    \"\"\"\n    index = []\n    # Regex: AIRS-CH0_signal_{obs}.parquet\n    pattern = re.compile(r'^AIRS-CH0_signal_(\\d+)\\.parquet$')\n    for file in files:\n        file_name = os.path.basename(file)\n        match = pattern.match(file_name)\n        if match:\n            planet_id = os.path.basename(os.path.dirname(file))\n            obs_num = int(match.group(1))\n            index.append((int(planet_id), obs_num))\n    # Sort by planet_id then obs_num\n    index.sort()\n    # Remove duplicates\n    index = list(dict.fromkeys(index))\n    if len(index) >= CHUNKS_SIZE and CHUNKS_SIZE > 0:\n        index_chunks = np.array_split(index, len(index)//CHUNKS_SIZE)\n    else:\n        index_chunks = [index]\n    return index_chunks\n\ndef bin_obs(arr, binning, axis=1):\n    \"\"\"\n    Bin an array along a specified axis by summing over non-overlapping bins.\n\n    Parameters:\n        arr (np.ndarray or np.ma.MaskedArray): Input array to bin.\n        binning (int): Bin size (number of elements per bin).\n        axis (int): Axis along which to bin (default: 1).\n\n    Returns:\n        np.ma.MaskedArray: Binned array with reduced size along the specified axis.\n    \"\"\"\n    bin_size = binning\n    arr = np.ma.masked_array(arr)\n    shape = list(arr.shape)\n    n_bins = shape[axis] // bin_size\n    new_shape = shape[:axis] + [n_bins, bin_size] + shape[axis+1:]\n    arr_reshaped = np.ma.reshape(arr, new_shape)\n    # Sum along the bin_size axis (axis=axis+1)\n    return np.ma.sum(arr_reshaped, axis=axis+1)\n\n\ndef median_filter_time(masked_arr, kernel_size=3):\n    \"\"\"\n    Apply a 1D median filter along the time axis (axis=1) for each batch.\n    Ignores masked voxels and preserves masked array structure.\n\n    Parameters:\n        masked_arr (np.ma.MaskedArray): Input array of shape (batch, time, X, Y).\n        kernel_size (int): Size of the median filter window (must be odd, default: 3).\n\n    Returns:\n        np.ma.MaskedArray: Median-filtered array of the same shape as input.\n    \"\"\"\n    assert kernel_size % 2 == 1, \"Kernel size must be odd!\"\n    batch_dim, time_dim, X, Y = masked_arr.shape\n    pad = kernel_size // 2\n    result = np.ma.masked_all(masked_arr.shape, dtype=masked_arr.dtype)\n    arr_data = masked_arr.data\n    arr_mask = masked_arr.mask\n\n    for b in range(batch_dim):\n        for t in range(time_dim):\n            lo = max(0, t - pad)\n            hi = min(time_dim, t + pad + 1)\n            window = arr_data[b, lo:hi, :, :]\n            window_mask = arr_mask[b, lo:hi, :, :]\n            window_ma = np.ma.masked_array(window, mask=window_mask)\n            # Compute median along the window (time) axis\n            median_vals = np.ma.median(window_ma, axis=0)\n            result.data[b, t] = median_vals.data\n            result.mask[b, t] = median_vals.mask\n\n    return result\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:45:12.324665Z","iopub.execute_input":"2025-09-30T18:45:12.324993Z","iopub.status.idle":"2025-09-30T18:45:12.339797Z","shell.execute_reply.started":"2025-09-30T18:45:12.324952Z","shell.execute_reply":"2025-09-30T18:45:12.338396Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Gather all signal file paths for training data\nfiles = glob.glob(os.path.join(path_folder, 'train','*','*'))\n#print(files[:5])  # Print first 5 files for inspection\nfor file in files[:5]:\n    file_name = file.split('/')[-1]\n    #print(file_name)  # Print file names for quick check\n\n# Get index chunks for batch processing (planet_id, obs_num pairs)\nindex_chunks = get_multiobs_index(files[:16], CHUNKS_SIZE)\n\n# Load calibration and axis info\ntrain_adc_info = pd.read_csv(os.path.join(path_folder, 'adc_info.csv'))\naxis_info = pd.read_parquet(os.path.join(path_folder,'axis_info.parquet'))\n\n# Flags to control cleaning steps\nDO_MASK = True      # Mask hot/dead pixels\nDO_THE_NL_CORR = False  # Apply non-linearity correction\nDO_DARK = True      # Subtract dark current\nDO_FLAT = True      # Apply flat field correction\nTIME_BINNING = True # Bin along time axis\nFILT = True         # Apply median filter\n\ncut_inf, cut_sup = 0, 356  # Wavelength axis cut\nl = cut_sup - cut_inf       # Length of wavelength axis\n\nfor index_chunk in index_chunks:\n    # Initialize masked arrays for cleaned AIRS and FGS1 data\n    AIRS_CH0_clean = np.ma.MaskedArray(np.zeros((CHUNKS_SIZE, 11250, 32, l)))\n    FGS1_clean = np.ma.MaskedArray(np.zeros((CHUNKS_SIZE, 135000, 32, 32)))\n    \n    for i in range(CHUNKS_SIZE):\n        # --- Load and calibrate AIRS-CH0 signal ---\n        df = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_signal_{index_chunk[i][1]}.parquet'))\n        signal = df.values.astype(np.float64).reshape((df.shape[0], 32, 356))\n        gain = train_adc_info['AIRS-CH0_adc_gain'][0]\n        offset = train_adc_info['AIRS-CH0_adc_offset'][0]\n        signal = ADC_convert(signal, gain, offset)\n        dt_airs = axis_info['AIRS-CH0-integration_time'].dropna().values\n        dt_airs[1::2] += 0.1  # Adjust integration time for odd frames\n        chopped_signal = signal[:, :, cut_inf:cut_sup]  # Crop wavelength axis\n        del signal, df\n        \n        # --- AIRS-CH0 cleaning steps ---\n        flat = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_calibration_{index_chunk[i][1]}/flat.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        dark = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_calibration_{index_chunk[i][1]}/dark.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        dead_airs = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_calibration_{index_chunk[i][1]}/dead.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        linear_corr = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_calibration_{index_chunk[i][1]}/linear_corr.parquet')).values.astype(np.float64).reshape((6, 32, 356))[:, :, cut_inf:cut_sup]\n        \n        if DO_MASK:\n            chopped_signal = mask_hot_dead(chopped_signal, dead_airs, dark)\n            AIRS_CH0_clean[i] = chopped_signal\n        else:\n            AIRS_CH0_clean[i] = chopped_signal\n        \n        if DO_THE_NL_CORR:\n            linear_corr_signal = apply_linear_corr(linear_corr, AIRS_CH0_clean[i])\n            AIRS_CH0_clean[i, :, :, :] = linear_corr_signal\n        del linear_corr\n        \n        if DO_DARK:\n            cleaned_signal = clean_dark(AIRS_CH0_clean[i], dead_airs, dark, dt_airs)\n            AIRS_CH0_clean[i] = cleaned_signal\n        del dark\n        \n        # --- Load and calibrate FGS1 signal ---\n        df = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_signal_{index_chunk[i][1]}.parquet'))\n        fgs_signal = df.values.astype(np.float64).reshape((df.shape[0], 32, 32))\n        FGS1_gain = train_adc_info['FGS1_adc_gain'][0]\n        FGS1_offset = train_adc_info['FGS1_adc_offset'][0]\n        fgs_signal = ADC_convert(fgs_signal, FGS1_gain, FGS1_offset)\n        dt_fgs1 = np.ones(len(fgs_signal)) * 0.1\n        dt_fgs1[1::2] += 0.1\n        chopped_FGS1 = fgs_signal\n        del fgs_signal, df\n        \n        # --- FGS1 cleaning steps ---\n        flat = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_calibration_{index_chunk[i][1]}/flat.parquet')).values.astype(np.float64).reshape((32, 32))\n        dark = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_calibration_{index_chunk[i][1]}/dark.parquet')).values.astype(np.float64).reshape((32, 32))\n        dead_fgs1 = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_calibration_{index_chunk[i][1]}/dead.parquet')).values.astype(np.float64).reshape((32, 32))\n        linear_corr = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_calibration_{index_chunk[i][1]}/linear_corr.parquet')).values.astype(np.float64).reshape((6, 32, 32))\n        \n        if DO_MASK:\n            chopped_FGS1 = mask_hot_dead(chopped_FGS1, dead_fgs1, dark)\n            FGS1_clean[i] = chopped_FGS1\n        else:\n            FGS1_clean[i] = chopped_FGS1\n        \n        if DO_THE_NL_CORR:\n            linear_corr_signal = apply_linear_corr(linear_corr, FGS1_clean[i])\n            FGS1_clean[i, :, :, :] = linear_corr_signal\n        del linear_corr\n        \n        if DO_DARK:\n            cleaned_signal = clean_dark(FGS1_clean[i], dead_fgs1, dark, dt_fgs1)\n            FGS1_clean[i] = cleaned_signal\n        del dark\n    \n    # --- Post-processing: CDS, filtering, binning ---\n    AIRS_cds = get_cds(AIRS_CH0_clean)\n    FGS1_cds = get_cds(FGS1_clean)\n    del AIRS_CH0_clean, FGS1_clean\n\n    if FILT:\n        AIRS_cds = median_filter_time(AIRS_cds)\n        FGS1_cds = median_filter_time(FGS1_cds)\n    \n    # Optional: bin along time axis to reduce data size\n    if TIME_BINNING:\n        AIRS_cds_binned = bin_obs(AIRS_cds, binning=1)\n        FGS1_cds_binned = bin_obs(FGS1_cds, binning=12)\n    else:\n        AIRS_cds_binned = AIRS_cds\n        FGS1_cds_binned = FGS1_cds\n    AIRS_cds_binned = AIRS_cds_binned.transpose(0, 1, 3, 2)\n    FGS1_cds_binned = FGS1_cds_binned.transpose(0, 1, 3, 2)\n    del AIRS_cds, FGS1_cds\n    \n    # --- Flat field correction ---\n    for i in range(CHUNKS_SIZE):\n        flat_airs = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/AIRS-CH0_calibration_{index_chunk[i][1]}/flat.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        flat_fgs = pd.read_parquet(os.path.join(path_folder, f'train/{index_chunk[i][0]}/FGS1_calibration_{index_chunk[i][1]}/flat.parquet')).values.astype(np.float64).reshape((32, 32))\n        if DO_FLAT:\n            corrected_AIRS_cds_binned = correct_flat_field(flat_airs, dead_airs, AIRS_cds_binned[i])\n            AIRS_cds_binned[i] = corrected_AIRS_cds_binned\n            corrected_FGS1_cds_binned = correct_flat_field(flat_fgs, dead_fgs1, FGS1_cds_binned[i])\n            FGS1_cds_binned[i] = corrected_FGS1_cds_binned\n    \n    AIRS_cds_binned = AIRS_cds_binned.transpose(0, 1, 3, 2)\n    FGS1_cds_binned = FGS1_cds_binned.transpose(0, 1, 3, 2)\n    \n    # --- Inpainting missing/bad pixels ---\n    # FGS1: inpaint along time axis as channels\n    data = FGS1_cds_binned[0, :, :, :].data         # shape: (time, x, y)\n    mask = FGS1_cds_binned[0, 0, :, :].mask         # shape: (x, y)\n    data = data.transpose(1, 2, 0)                 # shape: (x, y, time)\n    result_fgs1 = inpaint_biharmonic(data, mask, channel_axis=2)\n    result_fgs1 = result_fgs1.transpose(2, 0, 1)   # shape: (time, x, y)\n    \n    # AIRS: inpaint along wavelength axis as channels\n    data_airs = AIRS_cds_binned[0, :, :, :].data      # shape: (time, x, lambda)\n    mask_airs = AIRS_cds_binned[0, 0, :, :].mask      # shape: (x, lambda)\n    data_airs = data_airs.transpose(1, 2, 0)          # shape: (x, lambda, time)\n    result_airs = inpaint_biharmonic(data_airs, mask_airs, channel_axis=2)\n    result_airs = result_airs.transpose(2, 0, 1)\n    \n    # --- Save cleaned and inpainted data ---\n    chunk_name = '__'.join([f\"{pid}_{obs}\" for pid, obs in index_chunk])\n    np.savez(os.path.join(path_out, f'AIRS_clean_train_{chunk_name}.npz'), data=AIRS_cds_binned.data, mask=AIRS_cds_binned.mask)\n    np.savez(os.path.join(path_out, f'FGS1_clean_train_{chunk_name}.npz'), data=FGS1_cds_binned.data, mask=FGS1_cds_binned.mask)\n    np.savez(os.path.join(path_out, f'AIRS_cleaninp_train_{chunk_name}.npz'), data=result_airs, mask=AIRS_cds_binned.mask[0, :, :, :])\n    np.savez(os.path.join(path_out, f'FGS1_cleaninp_train_{chunk_name}.npz'), data=result_fgs1, mask=FGS1_cds_binned.mask[0, :, :, :])\n    \n    # Save index mapping as JSON for easy lookup\n    import json\n    index_chunk_serializable = [\n        [int(pid), int(obs)] if (not isinstance(pid, str)) else [str(pid), int(obs)]\n        for pid, obs in index_chunk\n    ]\n    with open(os.path.join(path_out, f'AIRS_clean_train_{chunk_name}_index.json'), 'w') as f_json:\n        json.dump(index_chunk_serializable, f_json)\n    \n    del AIRS_cds_binned\n    del FGS1_cds_binned","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T18:57:40.07323Z","iopub.execute_input":"2025-09-30T18:57:40.073614Z","iopub.status.idle":"2025-09-30T19:03:24.306574Z","shell.execute_reply.started":"2025-09-30T18:57:40.073579Z","shell.execute_reply":"2025-09-30T19:03:24.30543Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Read ADC info\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 356)\norig_imgs_cal = (orig_imgs / airs_gain) + airs_offset\n\nproc_path = '/kaggle/working/data_light_raw/AIRS_clean_train_1143471509_0.npz'\nproc_imgs = np.load(proc_path)  # shape: (num_frames, 32, l), l = cut_sup - cut_inf\nproc_imgs = np.ma.MaskedArray(proc_imgs['data'], proc_imgs['mask'])\n\nproc_imgs=proc_imgs[0,:,:,:]\ncut_inf, cut_sup = 0, 356     # as in your cleaning pipeline\n\n# Crop the original to compare matching region\norig_imgs_cropped = orig_imgs_cal[:, :, cut_inf:cut_sup]\n\nframes_to_plot = 9  # or any count you like\nplt.figure(figsize=(10, 22))\nfor idx in range(frames_to_plot):\n    # Plot original (calibrated, cropped) -- use odd frames only\n    orig_idx = 2 * idx + 1\n    plt.subplot(frames_to_plot, 2, 2*idx + 1)\n    plt.imshow(orig_imgs_cropped[orig_idx], cmap='gray', aspect='auto')\n    plt.title(f'Original Frame {orig_idx}')\n    plt.axis('off')\n    # Plot processed (already cleaned)\n    plt.subplot(frames_to_plot, 2, 2*idx + 2)\n    plt.imshow(proc_imgs[idx], cmap='gray', aspect='auto')\n    plt.title(f'Processed Frame {idx}')\n    plt.axis('off')\nplt.tight_layout()\nplt.show()\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:06:47.78483Z","iopub.execute_input":"2025-09-30T19:06:47.785252Z","iopub.status.idle":"2025-09-30T19:06:53.409005Z","shell.execute_reply.started":"2025-09-30T19:06:47.785219Z","shell.execute_reply":"2025-09-30T19:06:53.408068Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot compares the first nine odd-numbered time frames (frames 1, 3, 5, ..., 17) from the original, calibrated AIRS-CH0 detector images (left column) with their corresponding processed versions after cleaning and masking (right column) for a single planet observation.\n\nOriginal frames: Only odd-numbered frames are shown to match the binning-by-two performed during preprocessing, ensuring each processed frame aligns with its corresponding original frame in time.\n\nProcessed and masked frames: These are shown in sequence (frames 0 to 8), each corresponding to the original frame with index 1, 3, 5, etc.","metadata":{}},{"cell_type":"code","source":"# Read ADC info\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\ncut_inf, cut_sup = 0, 356  # adjust if needed as in preprocessing\n\n# Load original, processed, and inpainted data\n\n# --- Original ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 356)\norig_imgs_cal = (orig_imgs / airs_gain) + airs_offset\norig_imgs_cropped = orig_imgs_cal[:, :, cut_inf:cut_sup]\n\n# --- Processed (masked) ---\nproc_path = '/kaggle/working/data_light_raw/AIRS_clean_train_1143471509_0.npz'\nproc_npz = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_npz['data'], mask=proc_npz['mask'])\nproc_imgs = proc_imgs[0, :, :, :]  # shape: [frames, 32, cropped_wavelength]\n\n# --- Inpainted ---\ninpaint_path = '/kaggle/working/data_light_raw/AIRS_cleaninp_train_1143471509_0.npz'\ninp_npz = np.load(inpaint_path)\ninp_imgs = inp_npz['data']  # This is fully filled, not masked\ninp_mask = inp_npz['mask']  # Mask for reference if you want to display or highlight masked pixels\n\n# Select frame range for comparison\nframes_to_plot = 9  # or any count you like\nplt.figure(figsize=(12, 22))\nfor idx in range(frames_to_plot):\n    # Plot original (calibrated, cropped) -- use odd frames only\n    orig_idx = 2 * idx + 1\n    plt.subplot(frames_to_plot, 3, 3*idx + 1)\n    plt.imshow(orig_imgs_cropped[orig_idx], cmap='gray', aspect='auto')\n    plt.title(f'Original Frame {orig_idx}')\n    plt.axis('off')\n    \n    # Plot processed (masked)\n    plt.subplot(frames_to_plot, 3, 3*idx + 2)\n    plt.imshow(np.ma.filled(proc_imgs[idx], fill_value=np.nan), cmap='gray', aspect='auto')\n    plt.title(f'Processed Frame {idx}')\n    plt.axis('off')\n    \n    # Plot inpainted (filled)\n    plt.subplot(frames_to_plot, 3, 3*idx + 3)\n    plt.imshow(inp_imgs[idx], cmap='gray', aspect='auto')\n    plt.title(f'Inpainted Frame {idx}')\n    plt.axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:06:01.008622Z","iopub.execute_input":"2025-09-30T19:06:01.009032Z","iopub.status.idle":"2025-09-30T19:06:22.110788Z","shell.execute_reply.started":"2025-09-30T19:06:01.009Z","shell.execute_reply":"2025-09-30T19:06:22.109608Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot displays the first nine odd-numbered time frames (frames 1, 3, 5, ..., 17) for a single planet observation, comparing three stages of the AIRS-CH0 detector data:\n\nOriginal (left column): Calibrated raw images after gain and offset correction, showing only odd-numbered frames to match the binning-by-two performed during preprocessing. These frames may still contain hot/dead pixels and other artifacts.\n\nProcessed (middle column): Images after masking bad pixels and applying cleaning steps such as dark and flat field corrections. Masked regions appear as missing or blank areas, and each processed frame corresponds to its matching odd-numbered original frame.\n\nInpainted (right column): Images where missing or corrupted pixels have been filled using inpainting, resulting in fully continuous frames.","metadata":{}},{"cell_type":"code","source":"# --- Load ADC info for FGS1 ---\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nfgs1_gain = adc['FGS1_adc_gain'].iloc[0]\nfgs1_offset = adc['FGS1_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# --- ORIGINAL FGS1 SIGNAL ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/FGS1_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 32)\norig_imgs_cal = (orig_imgs / fgs1_gain) + fgs1_offset\n\n# --- PROCESSED FGS1 DATA ---\n# Use your saved processed file, e.g., \"/kaggle/working/data_light_raw/FGS1_clean_train_1143471509_0.npy\"\nproc_path = '/kaggle/working/data_light_raw/FGS1_clean_train_1143471509_0.npz'\nproc_imgs = np.load(proc_path)  # shape: (num_frames, 32, 32) for a single planet/obs. If shape is e.g., (batch, num_frames, 32, 32): use proc_imgs = proc_imgs[0]\nproc_imgs = np.ma.MaskedArray(proc_imgs['data'], proc_imgs['mask'])\nproc_imgs = proc_imgs[0,:,:,:]\nframes_to_plot = 9\nplt.figure(figsize=(10, 22))\nfor idx in range(frames_to_plot):\n    # Plot original (calibrated)\n    plt.subplot(frames_to_plot, 2, 2*idx + 1)\n    plt.imshow(orig_imgs_cal[idx], cmap='gray')\n    plt.title(f'Original Frame {idx}')\n    plt.axis('off')\n    # Plot processed\n    plt.subplot(frames_to_plot, 2, 2*idx + 2)\n    plt.imshow(proc_imgs[idx], cmap='gray')\n    plt.title(f'Processed Frame {idx}')\n    plt.axis('off')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:03:34.951913Z","iopub.execute_input":"2025-09-30T19:03:34.952365Z","iopub.status.idle":"2025-09-30T19:03:37.396457Z","shell.execute_reply.started":"2025-09-30T19:03:34.952328Z","shell.execute_reply":"2025-09-30T19:03:37.395402Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot displays the first nine time frames for a single planet observation, comparing the original calibrated FGS1 detector images (left column) with their corresponding processed versions (right column) after cleaning and masking.\n\nOriginal frames show the raw detector output after gain and offset calibration, but may still contain hot/dead pixels and other artifacts.\n\nProcessed frames have undergone masking of bad pixels and cleaning steps such as dark and flat field corrections, resulting in improved image quality and removal of corrupted regions.","metadata":{}},{"cell_type":"code","source":"# --- Load ADC info for FGS1 ---\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nfgs1_gain = adc['FGS1_adc_gain'].iloc[0]\nfgs1_offset = adc['FGS1_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# --- ORIGINAL FGS1 SIGNAL ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/FGS1_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 32)\norig_imgs_cal = (orig_imgs / fgs1_gain) + fgs1_offset\n\n# --- PROCESSED FGS1 DATA ---\nproc_path = '/kaggle/working/data_light_raw/FGS1_clean_train_1143471509_0.npz'\nproc_npz = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_npz['data'], mask=proc_npz['mask'])\nproc_imgs = proc_imgs[0,:,:,:]  # [num_frames, 32, 32]\n\n# --- INPAINTED FGS1 DATA ---\ninpaint_path = '/kaggle/working/data_light_raw/FGS1_cleaninp_train_1143471509_0.npz'\ninpaint_npz = np.load(inpaint_path)\ninpaint_imgs = inpaint_npz['data']  # shape: [num_frames, 32, 32]\ninpaint_mask = inpaint_npz['mask']  # [32, 32] or [num_frames, 32, 32] as reference\n\nframes_to_plot = 9\nplt.figure(figsize=(12, 22))\nfor idx in range(frames_to_plot):\n    # Plot original\n    plt.subplot(frames_to_plot, 3, 3*idx + 1)\n    plt.imshow(orig_imgs_cal[idx], cmap='gray')\n    plt.title(f'Original Frame {idx}')\n    plt.axis('off')\n\n    # Plot processed (masked)\n    plt.subplot(frames_to_plot, 3, 3*idx + 2)\n    plt.imshow(np.ma.filled(proc_imgs[idx], fill_value=np.nan), cmap='gray')\n    plt.title(f'Processed Frame {idx}')\n    plt.axis('off')\n\n    # Plot inpainted (filled)\n    plt.subplot(frames_to_plot, 3, 3*idx + 3)\n    plt.imshow(inpaint_imgs[idx], cmap='gray')\n    plt.title(f'Inpainted Frame {idx}')\n    plt.axis('off')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:03:37.397613Z","iopub.execute_input":"2025-09-30T19:03:37.398279Z","iopub.status.idle":"2025-09-30T19:03:44.426264Z","shell.execute_reply.started":"2025-09-30T19:03:37.398247Z","shell.execute_reply":"2025-09-30T19:03:44.425039Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot displays the first nine time frames for a single planet observation, comparing three stages of the FGS1 detector data:\n\nOriginal (left column): Calibrated raw images after gain and offset correction, which may still contain hot/dead pixels and other detector artifacts.\n\nProcessed (middle column): Images after masking bad pixels and applying cleaning steps such as dark and flat field corrections. Masked regions appear as missing or blank areas.\n\nInpainted (right column): Images where missing or corrupted pixels have been filled using inpainting, resulting in fully continuous frames.","metadata":{}},{"cell_type":"code","source":"# Load ADC info\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0]\n\n# Identify your planet and observation number\nplanet_id = 1143471509\nobs_num = 0\n\n# Load original data, calibrate, crop\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 356)\norig_imgs_cal = (orig_imgs / airs_gain) + airs_offset\ncut_inf, cut_sup = 0, 356\norig_imgs_cropped = orig_imgs_cal[:, :, cut_inf:cut_sup]\n\n# Load processed data: shape (num_frames, 32, l)\nproc_path = f'/kaggle/working/data_light_raw/AIRS_clean_train_{planet_id}_{obs_num}.npz'\nproc_imgs = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_imgs['data'], proc_imgs['mask'])\nif proc_imgs.ndim == 4:\n    proc_imgs = proc_imgs[0]  # Select first batch if there is a chunk dimension\n\n# Extract center slit (middle spatial row)\nmid_idx = orig_imgs_cropped.shape[1] // 2\norig_center = orig_imgs_cropped[:, mid_idx, :]  # Shape: (time, wavelength)\nproc_center = proc_imgs[:, mid_idx, :]          # Shape: (time, wavelength)\nprint(type(proc_center))\n# Plot original vs processed (time x wavelength images)\nplt.figure(figsize=(12, 6))\nplt.subplot(1, 2, 1)\nplt.imshow(orig_center.T, aspect='auto', cmap='gray')\nplt.title('Original Center-Slit (Time × Wavelength)')\nplt.xlabel('Time Frame')\nplt.ylabel('Wavelength Bin')\nplt.colorbar(label='Calibrated Flux')\n\nplt.subplot(1, 2, 2)\nplt.imshow(proc_center.T, aspect='auto', cmap='gray')\nplt.title('Processed Center-Slit (Time × Wavelength)')\nplt.xlabel('Time Frame')\nplt.ylabel('Wavelength Bin')\nplt.colorbar(label='Processed Flux')\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:03:44.427528Z","iopub.execute_input":"2025-09-30T19:03:44.427884Z","iopub.status.idle":"2025-09-30T19:03:48.171747Z","shell.execute_reply.started":"2025-09-30T19:03:44.427857Z","shell.execute_reply":"2025-09-30T19:03:48.170522Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot compares the time evolution of the center spatial row (slit) of the AIRS-CH0 detector for a single planet observation, before and after preprocessing:\n\nLeft panel (Original): Shows the calibrated, cropped center-slit data as a 2D image (time frame × wavelength bin) before cleaning. This view may contain detector artifacts, hot/dead pixels, and uncorrected noise.\n\nRight panel (Processed): Displays the same center-slit after cleaning steps such as masking, dark and flat field corrections, and binning. Masked or corrected regions are now handled, and the data is more suitable for scientific analysis.","metadata":{}},{"cell_type":"code","source":"# --- Load ADC info for FGS1 ---\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nfgs1_gain = adc['FGS1_adc_gain'].iloc[0]\nfgs1_offset = adc['FGS1_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# --- ORIGINAL FGS1 SIGNAL ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/FGS1_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 32)\norig_imgs_cal = (orig_imgs / fgs1_gain) + fgs1_offset\n\n# --- PROCESSED FGS1 DATA ---\nproc_path = '/kaggle/working/data_light_raw/FGS1_clean_train_1143471509_0.npz'\nproc_imgs = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_imgs['data'], proc_imgs['mask'])\nif proc_imgs.ndim == 4:\n    proc_imgs = proc_imgs[0]\n\n# Calculate white light curves\nlc_orig = orig_imgs_cal.sum(axis=(1,2))\nlc_proc = proc_imgs.sum(axis=(1,2))\n\n# Subsample original to match processed (odd and even indices)\nlc_orig_odd = lc_orig[1::2]\nlc_orig_even = lc_orig[0::2]\n\n# Normalize all curves\nlc_orig_odd_norm = lc_orig_odd / lc_orig_odd.mean()\nlc_orig_even_norm = lc_orig_even / lc_orig_even.mean()\nlc_proc_norm = lc_proc / lc_proc.mean()\n\nplt.figure(figsize=(12,6))\nplt.plot(lc_orig_even_norm, label='Original (even indices)', alpha=0.7)\nplt.plot(lc_orig_odd_norm, label='Original (odd indices)', alpha=0.7)\nplt.plot(lc_proc_norm, label='Processed', alpha=0.8, linewidth=2)\nplt.xlabel('Time (binned frame index)')\nplt.ylabel('Normalized flux in the frame')\nplt.title(f'FGS1 White Light Curve: Even vs Odd Original Indices and Processed\\nPlanet {planet_id}')\nplt.legend()\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:16:55.939199Z","iopub.execute_input":"2025-09-30T19:16:55.9396Z","iopub.status.idle":"2025-09-30T19:16:58.412443Z","shell.execute_reply.started":"2025-09-30T19:16:55.939574Z","shell.execute_reply":"2025-09-30T19:16:58.411279Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot compares the normalized white light curves (total flux summed over all pixels per frame) for a single planet observation using the FGS1 detector:\n\nOriginal (even and odd indices): The original calibrated frames are split into even and odd time indices to match the binning and frame selection performed during preprocessing. This allows for a direct comparison with the processed data, which typically results from combining or binning consecutive frames.\n\nProcessed: The cleaned and masked frames are shown as a single curve, representing the flux after all preprocessing steps (masking, corrections, binning).","metadata":{}},{"cell_type":"code","source":"# --- Load ADC info for FGS1 ---\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nfgs1_gain = adc['FGS1_adc_gain'].iloc[0]\nfgs1_offset = adc['FGS1_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# --- ORIGINAL FGS1 SIGNAL ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/FGS1_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 32)\norig_imgs_cal = (orig_imgs / fgs1_gain) + fgs1_offset\n\n# --- PROCESSED FGS1 DATA ---\nproc_path = '/kaggle/working/data_light_raw/FGS1_clean_train_1143471509_0.npz'\nproc_imgs = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_imgs['data'], proc_imgs['mask'])\nif proc_imgs.ndim == 4:\n    proc_imgs = proc_imgs[0]\n\n# Calculate white light curves\nlc_orig = orig_imgs_cal.sum(axis=(1,2))\nlc_proc = proc_imgs.sum(axis=(1,2))\n\n# Subsample the original to match processed (even and odd indices)\nlc_orig_even = lc_orig[0::2]\nlc_orig_odd = lc_orig[1::2]\n\n# Normalize all curves\nlc_orig_even_norm = lc_orig_even / lc_orig_even.mean()\nlc_orig_odd_norm = lc_orig_odd / lc_orig_odd.mean()\nlc_proc_norm = lc_proc / lc_proc.mean()\n\n# Plot three separate, aligned subplots\nfig, axes = plt.subplots(3, 1, figsize=(12, 9), sharex=True, sharey=True)\n\naxes[0].plot(lc_orig_even_norm, color='C0')\naxes[0].set_ylabel('Normalized Flux')\naxes[0].set_title('Original (even indices)')\n\naxes[1].plot(lc_orig_odd_norm, color='C1')\naxes[1].set_ylabel('Normalized Flux')\naxes[1].set_title('Original (odd indices)')\n\naxes[2].plot(lc_proc_norm, color='C2')\naxes[2].set_ylabel('Normalized Flux')\naxes[2].set_title('Processed (binned/cleaned)')\n\naxes[2].set_xlabel('Time (binned frame index)')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:17:33.289553Z","iopub.execute_input":"2025-09-30T19:17:33.290038Z","iopub.status.idle":"2025-09-30T19:17:35.6592Z","shell.execute_reply.started":"2025-09-30T19:17:33.289993Z","shell.execute_reply":"2025-09-30T19:17:35.658191Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot compares the normalized white light curves (total flux summed over all pixels per frame) for a single planet observation using the FGS1 detector, shown in three aligned subplots:\n\nOriginal (even indices): The top panel shows the white light curve for the even-numbered original frames.\n\nOriginal (odd indices): The middle panel shows the white light curve for the odd-numbered original frames.\n\nProcessed (binned/cleaned): The bottom panel shows the white light curve after cleaning and binning, which typically combines or averages consecutive frames.","metadata":{}},{"cell_type":"code","source":"# Compute absolute difference between consecutive points\njumps = np.abs(np.diff(lc_orig_even_norm))\n\n# Find index of largest jump\njump_idx = np.argmax(jumps)\n\nprint(f\"Largest jump between frames {jump_idx} and {jump_idx+1}\")\nprint(f\"Jump value: {lc_orig_even_norm[jump_idx+1] - lc_orig_even_norm[jump_idx]:.4f}\")\n\n# Optionally, plot with a marker\nplt.plot(lc_orig_even_norm)\nplt.axvline(jump_idx, color='red', linestyle='--', label='Largest jump')\nplt.xlabel('Frame Index')\nplt.ylabel('Normalized White Light Flux')\nplt.title('ORIGINAL EVEN White Light Curve With Largest Jump Marked')\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:19:17.472541Z","iopub.execute_input":"2025-09-30T19:19:17.474292Z","iopub.status.idle":"2025-09-30T19:19:17.772894Z","shell.execute_reply.started":"2025-09-30T19:19:17.47424Z","shell.execute_reply":"2025-09-30T19:19:17.771885Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This plot highlights the largest jump (outlier) in the normalized white light curve for the even-indexed original frames. The vertical red dashed line marks the frame where the most significant change in flux occurs between consecutive time points.","metadata":{}},{"cell_type":"code","source":"light_curve = lc_proc_norm\nprint(type(light_curve))\n\n# Compute absolute difference between consecutive points\njumps = np.abs(np.diff(light_curve))\n\n# Find index of largest jump\njump_idx = np.argmax(jumps)\n\nprint(f\"Largest jump between frames {jump_idx} and {jump_idx+1}\")\nprint(f\"Jump value: {light_curve[jump_idx+1] - light_curve[jump_idx]:.4f}\")\n\n# Optionally, plot with a marker\nplt.plot(light_curve)\nplt.axvline(jump_idx, color='red', linestyle='--', label='Largest jump')\nplt.xlabel('Frame Index')\nplt.ylabel('Normalized White Light Flux')\nplt.title('White Light Curve With Largest Jump Marked')\nplt.legend()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:20:55.637465Z","iopub.execute_input":"2025-09-30T19:20:55.638657Z","iopub.status.idle":"2025-09-30T19:20:55.881388Z","shell.execute_reply.started":"2025-09-30T19:20:55.638621Z","shell.execute_reply":"2025-09-30T19:20:55.880463Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We see in the processed signal the outlier is removed, but this isn't a guarantee and later we apply a median filter to remove any additional outlier values.","metadata":{}},{"cell_type":"code","source":"# --- Load ADC info for AIRS-CH0 ---\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# --- ORIGINAL AIRS-CH0 SIGNAL ---\norig_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_{obs_num}.parquet'\norig_df = pd.read_parquet(orig_path)\norig_imgs = orig_df.values.reshape(-1, 32, 356)\norig_imgs_cal = (orig_imgs / airs_gain) + airs_offset\n\n# --- PROCESSED AIRS-CH0 DATA ---\nproc_path = f'/kaggle/working/data_light_raw/AIRS_clean_train_{planet_id}_{obs_num}.npz'\nproc_imgs_npz = np.load(proc_path)\nproc_imgs = np.ma.MaskedArray(proc_imgs_npz['data'], proc_imgs_npz['mask'])\nif proc_imgs.ndim == 4:\n    proc_imgs = proc_imgs[0]\n\n# Optionally crop to match spectral region if necessary\ncut_inf, cut_sup = 39, 321  # Set to your analysis region\norig_imgs_cropped = orig_imgs_cal[:, :, cut_inf:cut_sup]\nproc_imgs_cropped = proc_imgs[:, :, :]  # If processed already cropped\n\n# Calculate white light curves\nlc_orig = orig_imgs_cropped.sum(axis=(1,2))\nlc_proc = proc_imgs_cropped.sum(axis=(1,2))\n\n# Subsample the original to match processed (even and odd indices)\nlc_orig_even = lc_orig[0::2]\nlc_orig_odd = lc_orig[1::2]\n\n# Normalize all curves\nlc_orig_even_norm = lc_orig_even / lc_orig_even.mean()\nlc_orig_odd_norm = lc_orig_odd / lc_orig_odd.mean()\nlc_proc_norm = lc_proc / lc_proc.mean()\n\n# Plot three separate, aligned subplots\nfig, axes = plt.subplots(3, 1, figsize=(12, 9), sharex=True, sharey=True)\n\naxes[0].plot(lc_orig_even_norm, color='C0')\naxes[0].set_ylabel('Normalized Flux')\naxes[0].set_title('AIRS Original (even indices)')\n\naxes[1].plot(lc_orig_odd_norm, color='C1')\naxes[1].set_ylabel('Normalized Flux')\naxes[1].set_title('AIRS Original (odd indices)')\n\naxes[2].plot(lc_proc_norm, color='C2')\naxes[2].set_ylabel('Normalized Flux')\naxes[2].set_title('AIRS Processed (binned/cleaned)')\n\naxes[2].set_xlabel('Time (binned frame index)')\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:22:47.17018Z","iopub.execute_input":"2025-09-30T19:22:47.170492Z","iopub.status.idle":"2025-09-30T19:22:50.88064Z","shell.execute_reply.started":"2025-09-30T19:22:47.170472Z","shell.execute_reply":"2025-09-30T19:22:50.879506Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This figure compares the normalized white light curves (total flux summed over all pixels per frame) for a single planet observation using the AIRS-CH0 spectrometer, shown in three aligned subplots:\n\nOriginal (even indices): The top panel shows the white light curve for the even-numbered original frames, which may include detector artifacts, cosmic ray hits, or instrumental noise.\n\nOriginal (odd indices): The middle panel shows the white light curve for the odd-numbered original frames, providing a complementary view of the raw data and helping to identify frame-to-frame variability or systematic effects.\n\nProcessed (binned/cleaned): The bottom panel shows the white light curve after cleaning and binning, which typically combines or averages consecutive frames and applies corrections for bad pixels, dark current, and flat field effects.","metadata":{}},{"cell_type":"code","source":"# -- LOAD ADC info --\nadc = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/adc_info.csv')\nairs_gain = adc['AIRS-CH0_adc_gain'].iloc[0]\nairs_offset = adc['AIRS-CH0_adc_offset'].iloc[0]\nfgs1_gain = adc['FGS1_adc_gain'].iloc[0]\nfgs1_offset = adc['FGS1_adc_offset'].iloc[0]\n\nplanet_id = 1143471509\nobs_num = 0\n\n# -- LOAD ORIGINAL AIRS-CH0 --\nairs_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/AIRS-CH0_signal_{obs_num}.parquet'\nairs_df = pd.read_parquet(airs_path)\norig_airs = airs_df.values.reshape(-1, 32, 356)\norig_airs_cal = (orig_airs / airs_gain) + airs_offset\ncut_inf, cut_sup = 39, 321  # Adjust as needed\norig_airs_crop = orig_airs_cal[:, :, cut_inf:cut_sup]\n\n# -- LOAD PROCESSED AIRS (masked .npz) --\nairs_proc_path = f'/kaggle/working/data_light_raw/AIRS_clean_train_{planet_id}_{obs_num}.npz'\nairs_proc_npz = np.load(airs_proc_path)\nairs_proc = np.ma.MaskedArray(airs_proc_npz['data'], airs_proc_npz['mask'])\nif airs_proc.ndim == 4:\n    airs_proc = airs_proc[0]\n\n# -- AIRS White Light Curves --\nlc_airs_orig = orig_airs_crop.sum(axis=(1,2))\nlc_airs_proc = airs_proc.sum(axis=(1,2))\nlc_airs_even = lc_airs_orig[0::2]\nlc_airs_odd = lc_airs_orig[1::2]\nlc_airs_even_norm = lc_airs_even / lc_airs_even.mean()\nlc_airs_odd_norm = lc_airs_odd / lc_airs_odd.mean()\nlc_airs_proc_norm = lc_airs_proc / lc_airs_proc.mean()\n\n# -- LOAD ORIGINAL FGS1 --\nfgs1_path = f'/kaggle/input/ariel-data-challenge-2025/train/{planet_id}/FGS1_signal_{obs_num}.parquet'\nfgs1_df = pd.read_parquet(fgs1_path)\norig_fgs1 = fgs1_df.values.reshape(-1, 32, 32)\norig_fgs1_cal = (orig_fgs1 / fgs1_gain) + fgs1_offset\n\n# -- LOAD PROCESSED FGS1 (masked .npz) --\nfgs1_proc_path = f'/kaggle/working/data_light_raw/FGS1_clean_train_{planet_id}_{obs_num}.npz'\nfgs1_proc_npz = np.load(fgs1_proc_path)\nfgs1_proc = np.ma.MaskedArray(fgs1_proc_npz['data'], fgs1_proc_npz['mask'])\nif fgs1_proc.ndim == 4:\n    fgs1_proc = fgs1_proc[0]\n\n# -- FGS1 White Light Curves --\nlc_fgs1_orig = orig_fgs1_cal.sum(axis=(1,2))\nlc_fgs1_proc = fgs1_proc.sum(axis=(1,2))\nlc_fgs1_even = lc_fgs1_orig[0::2]\nlc_fgs1_odd = lc_fgs1_orig[1::2]\nlc_fgs1_even_norm = lc_fgs1_even / lc_fgs1_even.mean()\nlc_fgs1_odd_norm = lc_fgs1_odd / lc_fgs1_odd.mean()\nlc_fgs1_proc_norm = lc_fgs1_proc / lc_fgs1_proc.mean()\n\n# -- PLOT 3x2 grid: left = AIRS, right = FGS1 --\nfig, axes = plt.subplots(3, 2, figsize=(14, 12), sharex='col', sharey='row')\n\n# AIRS plots\naxes[0,0].plot(lc_airs_even_norm, color='C0')\naxes[0,0].set_title('AIRS Original (even)')\naxes[0,0].set_ylabel('Normalized Flux')\n\naxes[1,0].plot(lc_airs_odd_norm, color='C1')\naxes[1,0].set_title('AIRS Original (odd)')\naxes[1,0].set_ylabel('Normalized Flux')\n\naxes[2,0].plot(lc_airs_proc_norm, color='C2')\naxes[2,0].set_title('AIRS Processed')\naxes[2,0].set_ylabel('Normalized Flux')\naxes[2,0].set_xlabel('Time (binned frame index)')\n\n# FGS1 plots\naxes[0,1].plot(lc_fgs1_even_norm, color='C0')\naxes[0,1].set_title('FGS1 Original (even)')\n\naxes[1,1].plot(lc_fgs1_odd_norm, color='C1')\naxes[1,1].set_title('FGS1 Original (odd)')\n\naxes[2,1].plot(lc_fgs1_proc_norm, color='C2')\naxes[2,1].set_title('FGS1 Processed')\naxes[2,1].set_xlabel('Time (binned frame index)')\n\n# Adjust layout\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:23:35.35695Z","iopub.execute_input":"2025-09-30T19:23:35.357338Z","iopub.status.idle":"2025-09-30T19:23:41.610574Z","shell.execute_reply.started":"2025-09-30T19:23:35.357309Z","shell.execute_reply":"2025-09-30T19:23:41.608859Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"This figure presents a 3x2 grid comparing the normalized white light curves (total flux summed over all pixels per frame) for a single planet observation, for both the AIRS-CH0 spectrometer (left column) and the FGS1 photometer (right column):\n\nTop row: White light curves for the even-numbered original frames, which may contain detector artifacts, cosmic ray hits, or instrumental noise.\n\nMiddle row: White light curves for the odd-numbered original frames, providing a complementary view of the raw data and helping to identify frame-to-frame variability or systematic effects.\n\nBottom row: White light curves after cleaning and binning (processed), which typically combine or average consecutive frames and apply corrections for bad pixels, dark current, and flat field effects.","metadata":{}},{"cell_type":"code","source":"lc_airs_proc_norm.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:04:02.632319Z","iopub.execute_input":"2025-09-30T19:04:02.632562Z","iopub.status.idle":"2025-09-30T19:04:02.639215Z","shell.execute_reply.started":"2025-09-30T19:04:02.632544Z","shell.execute_reply":"2025-09-30T19:04:02.638028Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lc_fgs1_proc_norm.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-30T19:04:02.640281Z","iopub.execute_input":"2025-09-30T19:04:02.640535Z","iopub.status.idle":"2025-09-30T19:04:02.660873Z","shell.execute_reply.started":"2025-09-30T19:04:02.640513Z","shell.execute_reply":"2025-09-30T19:04:02.659567Z"}},"outputs":[],"execution_count":null}]}