{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70367,"databundleVersionId":9188054,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport itertools\nimport os\nimport glob\nfrom astropy.stats import sigma_clip\n\nfrom tqdm import tqdm","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-22T05:22:20.850245Z","iopub.execute_input":"2024-09-22T05:22:20.850808Z","iopub.status.idle":"2024-09-22T05:22:21.849708Z","shell.execute_reply.started":"2024-09-22T05:22:20.850749Z","shell.execute_reply":"2024-09-22T05:22:21.848165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_folder = '/kaggle/input/ariel-data-challenge-2024/' # path to the folder containing the data\npath_out = '/kaggle/tmp/data_light_raw/' # path to the folder to store the light data\noutput_dir = '/kaggle/tmp/data_light_raw/' # path for the output directory","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:21.852312Z","iopub.execute_input":"2024-09-22T05:22:21.852984Z","iopub.status.idle":"2024-09-22T05:22:21.861999Z","shell.execute_reply.started":"2024-09-22T05:22:21.852928Z","shell.execute_reply":"2024-09-22T05:22:21.858127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if not os.path.exists(path_out):\n    os.makedirs(path_out)\n    print(f\"Directory {path_out} created.\")\nelse:\n    print(f\"Directory {path_out} already exists.\")","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:49.783727Z","iopub.execute_input":"2024-09-22T05:22:49.784181Z","iopub.status.idle":"2024-09-22T05:22:49.792146Z","shell.execute_reply.started":"2024-09-22T05:22:49.784138Z","shell.execute_reply":"2024-09-22T05:22:49.790844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding `train_adc_info.csv`\n\nThe **`train_adc_info.csv`** file contains **analog-to-digital conversion (ADC) parameters** for each observation (each exoplanet) in the training dataset. These parameters are crucial for restoring the original signal dynamic range from the digitized data.\n\n#### Column Descriptions:\n\n1. **planet_id**:\n   - Unique identifier for each exoplanet in the dataset, used to match observations of planetary transits.\n\n2. **FGS1_adc_offset**:\n   - This is the **offset** applied to the data from the Fine Guidance System 1 (FGS1) before digitization. It indicates how much the signal was shifted from the baseline before conversion. Use this value to **subtract** the baseline noise from the signal.\n\n3. **FGS1_adc_gain**:\n   - The **gain** applied to the FGS1 data before digitization. Gain refers to how much the signal was **amplified** before being converted to digital form. To restore the original data, you'll need to **divide** the signal by this gain.\n\n4. **AIRS-CH0_adc_offset**:\n   - Similar to FGS1, this is the **offset** applied to the data from the Ariel InfraRed Spectrometer Channel 0 (AIRS-CH0) before digitization. It needs to be subtracted to remove the bias introduced during digitization.\n\n5. **AIRS-CH0_adc_gain**:\n   - The **gain** applied to the AIRS-CH0 data before digitization. To retrieve the original signal, you must divide by this value, as the data was amplified before being digitized.\n\n6. **Star**:\n   - Identifies the **star** used in the exoplanet's simulation. Since the planet orbits this star, its light interacts with the exoplanet's atmosphere, influencing the spectrum you analyze.\n\n#### Why ADC Parameters are Important:\n\nThe ADC parameters are used to **calibrate the digitized data** back to its original form, ensuring the signal reflects the true observation:\n\n- **Offset** represents the base noise added during the digitization process. It needs to be **subtracted** from the signal.\n- **Gain** represents the amplification applied to the signal before digitization. It needs to be **divided** from the signal.\n\nThe equation for restoring the original signal from the digitized data is:\n\n\n**Original Signal** = $\\frac{\\mathbf{Digitized\\ Value} - \\mathbf{Offset}}{\\mathbf{Gain}}$\n\n\n#### Practical Use:\n\n- For **FGS1_signal.parquet** data, apply the **FGS1_adc_offset** and **FGS1_adc_gain** to correct the raw signals.\n- For **AIRS-CH0_signal.parquet** data, apply the **AIRS-CH0_adc_offset** and **AIRS-CH0_adc_gain** to calibrate the raw signals.\n\nThis step ensures that the data you're working with reflects the actual physical observations as closely as possible, allowing for more accurate analysis and machine learning modeling.\n","metadata":{}},{"cell_type":"code","source":"train_adc_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_adc_info.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:52.005924Z","iopub.execute_input":"2024-09-22T05:22:52.006384Z","iopub.status.idle":"2024-09-22T05:22:52.032978Z","shell.execute_reply.started":"2024-09-22T05:22:52.006339Z","shell.execute_reply":"2024-09-22T05:22:52.031896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_adc_info.sample(6)","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:52.623207Z","iopub.execute_input":"2024-09-22T05:22:52.623778Z","iopub.status.idle":"2024-09-22T05:22:52.660199Z","shell.execute_reply.started":"2024-09-22T05:22:52.623721Z","shell.execute_reply":"2024-09-22T05:22:52.65894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding `train_labels.csv`\n\nThe `train_labels` dataset contains the ground truth spectra for each exoplanet in the training set. Each row corresponds to an individual `planet_id` and its associated spectrum values at different wavelength intervals.\n\n- **planet_id**: A unique identifier for each planet in the dataset.\n- **wl_1, wl_2, ..., wl_278**: These columns represent the flux values of the spectrum at different wavelengths. Each column is labeled as `wl_X`, where X is the wavelength index. The flux values represent the exoplanet's transmission spectrum, which is the amount of starlight passing through the exoplanet's atmosphere at a specific wavelength.\n\nIn total, the dataset contains 284 columns:\n- The first column is the `planet_id`.\n- The next 283 columns (from `wl_1` to `wl_283`) contain the spectral data corresponding to each wavelength point.\n\nThis dataset will be used to train machine learning models to predict the spectra from the test dataset's exoplanet observations.","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:56.025702Z","iopub.execute_input":"2024-09-22T05:22:56.02664Z","iopub.status.idle":"2024-09-22T05:22:56.141502Z","shell.execute_reply.started":"2024-09-22T05:22:56.026587Z","shell.execute_reply":"2024-09-22T05:22:56.140177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:56.438843Z","iopub.execute_input":"2024-09-22T05:22:56.439265Z","iopub.status.idle":"2024-09-22T05:22:56.471072Z","shell.execute_reply.started":"2024-09-22T05:22:56.439225Z","shell.execute_reply":"2024-09-22T05:22:56.469817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding `wavelengths.csv`\n\nThe `wavelengths.csv` file contains the wavelength values corresponding to the flux data in the train and test signal files. Each column represents the central wavelength at which the flux values are measured.\n\n- **wl_1, wl_2, ..., wl_279**: These columns represent the wavelength values (in micrometers, µm). Each column corresponds to the wavelengths used in the spectral data. For example:\n  - `wl_1` = 0.705 µm (first wavelength point)\n  - `wl_2` = 1.951761 µm\n  - and so on, up to `wl_279`.\n\nThe data in this file will be used to associate each flux value in the signal data (both AIRS-CH0 and FGS1) with its corresponding wavelength. These wavelengths span the visible and infrared spectrum, capturing the range over which the instruments are sensitive.","metadata":{}},{"cell_type":"code","source":"wavelengths = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/wavelengths.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:57.25902Z","iopub.execute_input":"2024-09-22T05:22:57.259494Z","iopub.status.idle":"2024-09-22T05:22:57.280271Z","shell.execute_reply.started":"2024-09-22T05:22:57.259435Z","shell.execute_reply":"2024-09-22T05:22:57.27904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wavelengths.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:57.668052Z","iopub.execute_input":"2024-09-22T05:22:57.668548Z","iopub.status.idle":"2024-09-22T05:22:57.694139Z","shell.execute_reply.started":"2024-09-22T05:22:57.668499Z","shell.execute_reply":"2024-09-22T05:22:57.692689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding `axis_info.parquet`\n\nThe `axis_info.parquet` file contains important axis and integration time information for both instruments, AIRS-CH0 and FGS1, used in the Ariel simulation. This file is crucial for interpreting the structure of the time-series observations captured by these instruments during exoplanet transits.\n\nThe file includes the following columns:\n\n- **AIRS-CH0-axis0-h**: This column likely represents the time axis or the sequence of observations for the AIRS-CH0 instrument in units of hours (`h`). Each row corresponds to a specific time at which a measurement was taken.\n  \n- **AIRS-CH0-axis2-um**: This column represents the wavelength axis for the AIRS-CH0 instrument in units of micrometers (`um`). Each row specifies the wavelength value corresponding to the measurements captured at that time.\n\n- **AIRS-CH0-integration_time**: This column indicates the integration time for the AIRS-CH0 instrument, measured in seconds. The integration time reflects how long the instrument observed for each data point. It alternates between 0.1 seconds and 4.5 seconds, allowing for different temporal resolutions during the observations.\n\n- **FGS1-axis0-h**: Similar to AIRS-CH0-axis0-h, this column represents the time axis for the FGS1 instrument in units of hours (`h`). Each row specifies the time when the FGS1 observations were made.\n\nThis data is important for understanding the exact timing and wavelength of each observation, which is necessary for correctly calibrating and interpreting the observed signals from the exoplanet transits.\n\n#### Example:\nFor instance, the first row in the dataset shows:\n- **AIRS-CH0-axis0-h**: `0.000028` (hours)\n- **AIRS-CH0-axis2-um**: `4.078463` (micrometers)\n- **AIRS-CH0-integration_time**: `0.1` (seconds)\n- **FGS1-axis0-h**: `0.000028` (hours)\n\nThis indicates that at a time of 0.000028 hours, the AIRS-CH0 instrument measured at a wavelength of 4.078463 micrometers with an integration time of 0.1 seconds, and the FGS1 instrument made an observation at the same time.\n","metadata":{}},{"cell_type":"code","source":"axis_info = pd.read_parquet('/kaggle/input/ariel-data-challenge-2024/axis_info.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:58.240874Z","iopub.execute_input":"2024-09-22T05:22:58.241984Z","iopub.status.idle":"2024-09-22T05:22:58.41121Z","shell.execute_reply.started":"2024-09-22T05:22:58.24192Z","shell.execute_reply":"2024-09-22T05:22:58.40983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"axis_info.shape","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:59.092083Z","iopub.execute_input":"2024-09-22T05:22:59.092553Z","iopub.status.idle":"2024-09-22T05:22:59.100718Z","shell.execute_reply.started":"2024-09-22T05:22:59.092505Z","shell.execute_reply":"2024-09-22T05:22:59.099265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"axis_info.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:22:59.618338Z","iopub.execute_input":"2024-09-22T05:22:59.619361Z","iopub.status.idle":"2024-09-22T05:22:59.633748Z","shell.execute_reply.started":"2024-09-22T05:22:59.619312Z","shell.execute_reply":"2024-09-22T05:22:59.632539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_adc_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/test_adc_info.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:24:36.850653Z","iopub.execute_input":"2024-09-22T05:24:36.851787Z","iopub.status.idle":"2024-09-22T05:24:36.863979Z","shell.execute_reply.started":"2024-09-22T05:24:36.851735Z","shell.execute_reply":"2024-09-22T05:24:36.862434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_adc_info.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:24:47.90415Z","iopub.execute_input":"2024-09-22T05:24:47.904632Z","iopub.status.idle":"2024-09-22T05:24:47.918503Z","shell.execute_reply.started":"2024-09-22T05:24:47.904588Z","shell.execute_reply":"2024-09-22T05:24:47.917303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:27:35.96726Z","iopub.execute_input":"2024-09-22T05:27:35.967811Z","iopub.status.idle":"2024-09-22T05:27:36.002057Z","shell.execute_reply.started":"2024-09-22T05:27:35.96776Z","shell.execute_reply":"2024-09-22T05:27:36.000576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:27:45.792953Z","iopub.execute_input":"2024-09-22T05:27:45.793435Z","iopub.status.idle":"2024-09-22T05:27:45.821781Z","shell.execute_reply.started":"2024-09-22T05:27:45.79338Z","shell.execute_reply":"2024-09-22T05:27:45.820318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training Data Structure\n\nThe training dataset is organized in a hierarchical structure, with each exoplanet observation stored in its own directory. The folder structure is organized as follows:\n","metadata":{}},{"cell_type":"markdown","source":"\n#### Description of the Train Folder Structure:\n- **train/**: This is the root folder containing the training data. Each subfolder corresponds to a unique exoplanet identified by its `planet_id`.\n  \n- **[planet_id]/**: Each folder under `train/` represents the observational data for a specific exoplanet. The `planet_id` is a unique identifier for each planet.\n  \n  - **AIRS-CH0_calibration/**: This folder contains calibration data for the AIRS-CH0 instrument, which is used to correct the signals from the observations. Each calibration file is essential to reduce noise and improve the accuracy of the data.\n    \n    - **dark.parquet**: Captures the dark frame data to subtract sensor bias and thermal noise.\n    - **dead.parquet**: Identifies dead or hot pixels on the sensor that either produce no response or a constant high signal.\n    - **flat.parquet**: Used to correct pixel-to-pixel sensitivity variations.\n    - **linear_corr.parquet**: Provides correction for sensor non-linearities.\n    - **read.parquet**: Contains the read noise data to account for electronic noise during data capture.\n\n  - **FGS1_calibration/**: Similar to the AIRS-CH0_calibration folder, this folder contains calibration data for the FGS1 instrument, including the same types of files: dark, dead, flat, linear_corr, and read.\n\n  - **AIRS-CH0_signal.parquet**: This file contains the actual observation data from the AIRS-CH0 instrument after it has passed through the sensor.\n\n  - **FGS1_signal.parquet**: Contains the observational signal data from the FGS1 instrument.\n\nEach exoplanet's folder includes both the raw signal data and the corresponding calibration data for both instruments (AIRS-CH0 and FGS1). This structure allows for systematic pre-processing of the data using the provided calibration frames.","metadata":{}},{"cell_type":"markdown","source":"### Understanding `AIRS-CH0_signal.parquet`\n\nThe `AIRS-CH0_signal.parquet` file contains the raw observational data collected by the AIRS-CH0 instrument during the exoplanet transit. This file stores a large number of sequential 2D image frames, where each column represents a specific pixel from the instrument's spectral focal plane over time. The rows correspond to time steps in the observation period, capturing the signal variations during the transit.\n\n#### Data Structure:\n\n- **Rows (Time steps)**: Each row corresponds to a distinct time step in the observation. For instance, a time step represents a frame of the star and exoplanet during the transit.\n  \n- **Columns (Pixel values)**: The columns represent the pixel intensity values from the 2D spectral image captured by AIRS-CH0. There are **11,392 columns**, each associated with a different pixel in the image. These values indicate the raw intensity measurements, which are likely to include noise and other artifacts.\n","metadata":{}},{"cell_type":"code","source":"AIRS_CHO_Signal = pd.read_parquet('/kaggle/input/ariel-data-challenge-2024/train/100468857/AIRS-CH0_signal.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:36:52.106656Z","iopub.execute_input":"2024-09-22T05:36:52.107159Z","iopub.status.idle":"2024-09-22T05:36:54.816574Z","shell.execute_reply.started":"2024-09-22T05:36:52.107113Z","shell.execute_reply":"2024-09-22T05:36:54.815228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AIRS_CHO_Signal.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:36:56.111064Z","iopub.execute_input":"2024-09-22T05:36:56.111544Z","iopub.status.idle":"2024-09-22T05:36:56.138171Z","shell.execute_reply.started":"2024-09-22T05:36:56.111489Z","shell.execute_reply":"2024-09-22T05:36:56.136795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AIRS_CHO_Signal.shape","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:47:45.413104Z","iopub.execute_input":"2024-09-22T05:47:45.413643Z","iopub.status.idle":"2024-09-22T05:47:45.422678Z","shell.execute_reply.started":"2024-09-22T05:47:45.413595Z","shell.execute_reply":"2024-09-22T05:47:45.421276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Understanding `FGS1_signal.parquet`\n\nThe `FGS1_Signal.parquet` file contains the raw signal data collected by the Fine Guidance Sensor 1 (FGS1) during the exoplanet transit observation. FGS1 is an essential part of the Ariel mission, primarily used for precise pointing and control of the telescope. This data might also carry information about the star's brightness, which could help in tracking the transit and extracting valuable information about the exoplanet.\n\n#### Data Structure:\n\n- **Rows (Time steps)**: Each row represents a single time step in the observation period. As with the AIRS-CH0 data, these time steps capture the star's and exoplanet's signal during the transit.\n  \n- **Columns (Pixel values)**: The columns correspond to pixel intensity values from the FGS1 sensor. There are **1,024 columns**, each representing a different pixel. These values indicate the brightness variations measured by the FGS1 sensor at each time step.\n\n\n#### Data Insights:\n\n- The **FGS1** sensor primarily monitors the star's brightness to maintain the telescope's pointing accuracy. However, the sensor also captures small brightness fluctuations that can help confirm the transit signal.\n- The pixel intensity values recorded in this dataset will vary due to changes in the star's light curve as the planet transits across the star.\n- Like the AIRS-CH0 dataset, the FGS1 data requires calibration to remove noise and sensor imperfections. Calibration files for FGS1 (dark, flat, linear_corr, etc.) will help in this process.\n","metadata":{}},{"cell_type":"code","source":"FGS1_Signal = pd.read_parquet('/kaggle/input/ariel-data-challenge-2024/train/100468857/FGS1_signal.parquet')","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:47:46.717899Z","iopub.execute_input":"2024-09-22T05:47:46.718353Z","iopub.status.idle":"2024-09-22T05:47:48.279212Z","shell.execute_reply.started":"2024-09-22T05:47:46.718312Z","shell.execute_reply":"2024-09-22T05:47:48.277791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FGS1_Signal.head()","metadata":{"execution":{"iopub.status.busy":"2024-09-22T05:47:48.281087Z","iopub.execute_input":"2024-09-22T05:47:48.281893Z","iopub.status.idle":"2024-09-22T05:47:48.303274Z","shell.execute_reply.started":"2024-09-22T05:47:48.281846Z","shell.execute_reply":"2024-09-22T05:47:48.301949Z"},"trusted":true},"execution_count":null,"outputs":[]}]}