{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":12846694,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n### Understanding the Challenge\n\n\n*   **The Goal:** We need to analyze simulated time-series data from two telescope instruments (AIRS-CH0 and FGS1) to extract the \"spectrum\" of an exoplanet's atmosphere. A spectrum tells us how much starlight the atmosphere blocks at different wavelengths, which reveals its chemical composition.\n*   **The Data:** The raw data is a series of images (frames) over time, showing the starlight. When the planet transits (passes in front of its star), the light dims slightly. This dimming is what we need to measure. The data is noisy and contains instrumental effects we need to correct.\n*   **The Output:** For each planet in the test set, we must predict its spectrum (a mean value, `μ_user`, for each wavelength) and the uncertainty of our prediction (`σ_user`).\n*   **The Evaluation:** The score is based on the Gaussian Log-likelihood (GLL). This metric heavily penalizes predictions that are confidently wrong (i.e., the ground truth is far from your predicted mean, and your predicted uncertainty is small). Therefore, estimating a reasonable uncertainty is just as important as predicting the spectrum itself.\n\n### Step 1: Loading and Exploring Metadata\n\nThe raw signal data is huge (over 260 GB!), so we should never try to load it all at once. The best place to start is with the smaller metadata files. They will give us a high-level overview of the dataset structure.\n\nLet's begin by loading `train.csv`, `train_star_info.csv`, `wavelengths.csv`, `adc_info.csv`, and `axis_info.parquet`. We'll also merge the training information to get a complete picture for each planet.\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom pathlib import Path\n\n# --- 1. Define the base path for the data ---\n# This is the standard path in a Kaggle environment\nBASE_PATH = Path(\"/kaggle/input/ariel-data-challenge-2025\")\n\n# --- 2. Load all the main metadata files ---\nprint(\"Loading all top-level metadata files...\")\ntry:\n    train_df = pd.read_csv(BASE_PATH / \"train.csv\")\n    train_star_info_df = pd.read_csv(BASE_PATH / \"train_star_info.csv\")\n    wavelengths_df = pd.read_csv(BASE_PATH / \"wavelengths.csv\")\n    adc_info_df = pd.read_csv(BASE_PATH / \"adc_info.csv\")\n    axis_info_df = pd.read_parquet(BASE_PATH / \"axis_info.parquet\")\n    sample_submission_df = pd.read_csv(BASE_PATH / \"sample_submission.csv\")\n    print(\"All files loaded successfully!\")\nexcept FileNotFoundError as e:\n    print(f\"Error loading files: {e}\")\n    print(\"Please ensure your notebook is connected to the competition dataset.\")\n\n# --- 3. Display the head of each dataframe ---\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"1. Ground Truth Spectra (train.csv)\")\nprint(\"=\"*50)\n# This file contains the target values (the true spectra) we need to predict.\n# The columns are likely numbered 0 to 282, corresponding to different wavelengths.\ndisplay(train_df.head())\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"2. Star and Planet Physical Info (train_star_info.csv)\")\nprint(\"=\"*50)\n# These are the physical characteristics of each system. These will be our primary features.\ndisplay(train_star_info_df.head())\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"3. Wavelength Grid (wavelengths.csv)\")\nprint(\"=\"*50)\n# This file maps the column indices from train.csv to physical wavelengths (in µm).\n# It's the 'x-axis' for our spectra.\ndisplay(wavelengths_df.head())\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"4. ADC Conversion Parameters (adc_info.csv)\")\nprint(\"=\"*50)\n# These are the gain and offset values to restore the raw signal data's dynamic range.\ndisplay(adc_info_df.head())\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"5. Axis Information (axis_info.parquet)\")\nprint(\"=\"*50)\n# This should contain timing or coordinate information for the raw signal data.\ndisplay(axis_info_df.head())\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"6. Sample Submission Format (sample_submission.csv)\")\nprint(\"=\"*50)\n# This shows us exactly how our final output file must be structured.\ndisplay(sample_submission_df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:16:29.662611Z","iopub.execute_input":"2025-06-27T20:16:29.662859Z","iopub.status.idle":"2025-06-27T20:16:30.287618Z","shell.execute_reply.started":"2025-06-27T20:16:29.662838Z","shell.execute_reply":"2025-06-27T20:16:30.286618Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n### Step 2: Combining and Reshaping Data for Analysis\n\nLet's perform two crucial actions to make this data easier to work with:\n\n1.  **Merge** `train.csv` and `train_star_info.csv` into a single, comprehensive dataframe.\n2.  **Reshape** `wavelengths.csv` into a more usable format and add the corresponding instrument name for each wavelength.\n\n","metadata":{}},{"cell_type":"code","source":"# --- 1. Merge the training data ---\n# We'll use an 'inner' merge, which is the default. This ensures that we only keep\n# planets that appear in both files.\nmerged_train_df = pd.merge(train_df, train_star_info_df, on=\"planet_id\")\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"1. Merged Training Data (Spectra + Star/Planet Info)\")\nprint(\"=\"*50)\nprint(f\"Shape of merged data: {merged_train_df.shape}\")\nprint(\"This dataframe now contains our primary features (star/planet info) and targets (spectra).\")\ndisplay(merged_train_df.head())\n\n\n# --- 2. Reshape the wavelengths data ---\n# The original format is wide (1 row, 283 columns). Let's make it long.\nwavelengths_long_df = wavelengths_df.T.reset_index()\nwavelengths_long_df.columns = ['wl_id', 'wavelength_um']\n\n# Now, let's add the instrument based on the wavelength, as described in the report.\n# FGS1 is a single point at 0.7 um. The rest are AIRS-CH0.\nwavelengths_long_df['instrument'] = np.where(wavelengths_long_df['wavelength_um'] < 1.0, 'FGS1', 'AIRS-CH0')\n\nprint(\"\\n\" + \"=\"*50)\nprint(\"2. Reshaped Wavelengths Data\")\nprint(\"=\"*50)\nprint(f\"Shape of reshaped wavelength data: {wavelengths_long_df.shape}\")\nprint(\"This format is much easier to use for plotting and lookups.\")\ndisplay(wavelengths_long_df.head()) # Shows the FGS1 point\ndisplay(wavelengths_long_df.tail()) # Shows the AIRS-CH0 points\n\n# Let's double-check the counts per instrument\nprint(\"\\nInstrument counts from reshaped wavelength data:\")\nprint(wavelengths_long_df['instrument'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:16:30.289903Z","iopub.execute_input":"2025-06-27T20:16:30.290184Z","iopub.status.idle":"2025-06-27T20:16:30.334043Z","shell.execute_reply.started":"2025-06-27T20:16:30.290162Z","shell.execute_reply":"2025-06-27T20:16:30.333151Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n*   We have a `merged_train_df` with 1100 planets, where each row contains both the physical properties (features) and the full ground truth spectrum (targets).\n*   We have a `wavelengths_long_df` that acts as a handy lookup table for our spectral data points.\n\nNow, let's dive in and visualize this data.\n\n### Step 3: Visualizing Feature Distributions and a Sample Spectrum\n\nWe'll do two things in this step:\n\n1.  **Analyze the Features:** We will create histograms for the physical parameters of the star-planet systems (`Rs`, `Ts`, `Mp`, etc.). This helps us understand the diversity of our training set. Are we dealing with a wide range of star types and planet sizes? Are there any unusual outliers?\n2.  **Analyze the Target:** We will pick a single planet and plot its ground-truth spectrum. This will give us our first look at the signal we are ultimately trying to predict. What do these spectra look like? Are there clear absorption features?\n\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport warnings\n\n# Suppress the specific FutureWarning from Seaborn\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\n# Set a nice theme for Seaborn plots\nsns.set_theme(style=\"whitegrid\")\n\n# These are the columns we want to inspect\nfeature_cols = ['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']\n\npalette = sns.color_palette(\"viridis\", len(feature_cols))\n\nfig, axes = plt.subplots(nrows=2, ncols=4, figsize=(16, 8))\naxes = axes.flatten()\n\nfor i, col in enumerate(feature_cols):\n    ax = axes[i]\n    sns.histplot(\n        data=merged_train_df,\n        x=col,\n        ax=ax,\n        kde=True,\n        color=palette[i] # Assign the i-th color from our new palette\n    )\n    ax.set_title(f'Distribution of {col}', fontsize=14)\n    ax.set_yscale('log')\n    ax.set_xlabel('')\n\nif len(feature_cols) < len(axes):\n    for i in range(len(feature_cols), len(axes)):\n        axes[i].set_visible(False)\n\nfig.suptitle('Distributions of Star and Planet Physical Parameters', fontsize=20, y=1.02)\nplt.tight_layout(rect=[0, 0, 1, 0.98])\nplt.show()\n\n\n# ==============================================================================\n# PLOT 2: Spectrum of a Single Planet with custom colors\n# ==============================================================================\n\nprint(\"\\n--- 2. Visualizing a Single Spectrum with custom colors ---\")\n\n# --- Data Preparation (same as before) ---\nplanet_index = 0\nplanet_sample = merged_train_df.iloc[[planet_index]]\nplanet_id_sample = planet_sample['planet_id'].values[0]\nwl_cols = [col for col in train_df.columns if col.startswith('wl_')]\nspectrum_sample_long = planet_sample.melt(\n    id_vars=['planet_id'], value_vars=wl_cols,\n    var_name='wl_id', value_name='spectrum_value'\n)\nspectrum_to_plot = pd.merge(spectrum_sample_long, wavelengths_long_df, on='wl_id')\n\n\ninstrument_colors = {\n    \"FGS1\": \"#0077b6\",  # A strong, clear blue\n    \"AIRS-CH0\": \"#d62728\" # A classic, bold red\n}\n\n# --- Plotting with Seaborn ---\nplt.figure(figsize=(14, 7))\n\nsns.lineplot(\n    data=spectrum_to_plot,\n    x='wavelength_um',\n    y='spectrum_value',\n    hue='instrument',\n    style='instrument',\n    markers=True,\n    dashes=False,\n    palette=instrument_colors # Apply our custom color dictionary\n)\n\nplt.title(f'Ground Truth Spectrum for Planet ID: {planet_id_sample}', fontsize=18)\nplt.xlabel('Wavelength (µm)', fontsize=12)\nplt.ylabel('Transit Depth (unitless)', fontsize=12)\nplt.legend(title='Instrument', fontsize=11)\nplt.grid(True, which='both', linestyle='--', linewidth=0.5)\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:24:36.135438Z","iopub.execute_input":"2025-06-27T20:24:36.136511Z","iopub.status.idle":"2025-06-27T20:24:40.599603Z","shell.execute_reply.started":"2025-06-27T20:24:36.13647Z","shell.execute_reply":"2025-06-27T20:24:40.598511Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Analysis of our Plots\n\n**1. Distributions of Star and Planet Physical Parameters**\n\n\n*   **Rs (Stellar Radius):** Most stars have a radius between 0.8 and 1.5 times that of our Sun. It's a right-skewed distribution.\n*   **Ms (Stellar Mass):** Similar to radius, most stars are around the mass of our Sun (1.0), with a distribution skewed towards slightly more massive stars.\n*   **Ts (Stellar Temperature):** A nice, roughly normal distribution centered around 5800K, which is very similar to our Sun. This means we're mostly dealing with Sun-like stars.\n*   **Mp (Planetary Mass):** This is heavily right-skewed. Most planets in the dataset are \"small\" (less than ~2 Earth masses), but there's a long tail of much more massive planets. This could be important for modeling.\n*   **P (Orbital Period):** Appears to be a fairly uniform distribution between ~3 and ~7 days. These are all \"hot\" planets, orbiting very close to their stars.\n*   **sma (Semi-major Axis):** A distribution centered around 10-12 stellar radii. This is the orbital distance.\n*   **i (Inclination):** Tightly clustered between 87 and 90 degrees. This is exactly what we'd expect. An inclination of 90 degrees is a perfect edge-on orbit. We can only observe a transit if the inclination is very close to 90.\n\n**2. Ground Truth Spectrum for Planet ID: 34983**\n\nThis plot is perfect and gives us our first real look at the prize!\n\n*   **This is the Signal:** This is what we are trying to recover. The y-axis shows the \"transit depth,\" which is the fraction of the star's light blocked by the planet.\n*   **The \"Wiggles\" are Everything:** The plot is not a flat line. The bumps and dips, especially in the red AIRS-CH0 data, are the chemical fingerprints of the planet's atmosphere. For example, the big bump around 3.3 µm and the dip near 2.7 µm indicate that the atmosphere absorbs light differently at those wavelengths. This is the information we need to extract from the noisy raw data.\n*   **Instrument Separation:** We can clearly see the single FGS1 data point (blue) at 0.7 µm is distinct from the detailed AIRS-CH0 spectrum (red) which starts at ~1.95 µm.\n\n### Step 4: Fixing the Histograms and Diving into the Raw Signal\n\nLet's do two things now:\n1.  Correct our histogram plot to properly visualize the feature distributions.\n2.  Take the next logical step: load the **raw time-series signal data** for this *exact same planet* (ID 34983) and create a \"light curve\" – a plot of brightness over time. This will show us the transit event that the spectrum was derived from.\n\n","metadata":{}},{"cell_type":"code","source":"\n\nplanet_id_sample = 34983\nfgs1_path = BASE_PATH / f\"train/{planet_id_sample}/FGS1_signal_0.parquet\"\nairs_path = BASE_PATH / f\"train/{planet_id_sample}/AIRS-CH0_signal_0.parquet\"\n\nprint(f\"Loading raw signal for planet {planet_id_sample}...\")\nfgs1_raw_df = pd.read_parquet(fgs1_path)\nairs_raw_df = pd.read_parquet(airs_path)\n\ngain = adc_info_df['FGS1_adc_gain'].iloc[0]\noffset = adc_info_df['FGS1_adc_offset'].iloc[0]\n\nfgs1_corrected = fgs1_raw_df * gain + offset\nairs_corrected = airs_raw_df * gain + offset\n\nfgs1_light_curve = fgs1_corrected.sum(axis=1)\nairs_light_curve = airs_corrected.sum(axis=1)\n\n\n\nprint(\"Plotting light curves...\")\n\n\nfig, axes = plt.subplots(nrows=2, ncols=1, figsize=(14, 9))\n\n# FGS1 Plot (Top)\nsns.lineplot(x=fgs1_light_curve.index, y=fgs1_light_curve, ax=axes[0], color=\"#0077b6\", lw=0.5)\naxes[0].set_title('FGS1 Light Curve', fontsize=16)\naxes[0].set_ylabel('Total Flux (Arbitrary Units)', fontsize=12)\n# Add an x-label to the top plot for clarity\naxes[0].set_xlabel('Frame Number (Time)', fontsize=12)\n\n\n# AIRS-CH0 Plot (Bottom)\nsns.lineplot(x=airs_light_curve.index, y=airs_light_curve, ax=axes[1], color=\"#d62728\", lw=0.5)\naxes[1].set_title('AIRS-CH0 Light Curve', fontsize=16)\naxes[1].set_ylabel('Total Flux (Arbitrary Units)', fontsize=12)\naxes[1].set_xlabel('Frame Number (Time)', fontsize=12)\n\nfig.suptitle(f\"Raw Light Curves for Planet ID: {planet_id_sample}\", fontsize=20, y=1.01)\nplt.tight_layout(rect=[0, 0, 1, 0.97])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:38:51.790107Z","iopub.execute_input":"2025-06-27T20:38:51.790476Z","iopub.status.idle":"2025-06-27T20:38:58.991096Z","shell.execute_reply.started":"2025-06-27T20:38:51.790447Z","shell.execute_reply":"2025-06-27T20:38:58.990199Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n\n\n### Analysis of the Raw Light Curves\n\nThis is the most critical plot so far. It reveals the core challenge of this competition.\n\n*   **Where is the Transit?** We expected to see a clear \"U\" or \"V\" shaped dip, but instead, we see a thick, noisy band of data. The transit signal (the dip) is completely buried in the noise! The change in brightness due to the planet is much, much smaller than the random noise and other instrumental effects.\n*   **Instrumental Effects:** The description mentioned that the flux levels oscillate between low and high counts due to the \"up-the-ramp\" sampling. This high-frequency oscillation is likely the cause of the very thick band we see in the plot. Our simple `sum()` is not enough to de-noise this.\n*   **Spikes:** The FGS1 light curve shows a few prominent spikes. These are likely cosmic ray hits or other artifacts that need to be cleaned.\n*   **The Real Task:** This plot clarifies our mission. The goal is not just to find a dip, but to perform sophisticated signal processing to:\n    1.  Remove instrumental noise (like the up-the-ramp effect).\n    2.  Correct for artifacts (like cosmic rays).\n    3.  Average the signal in a way that the tiny transit dip becomes visible.\n    4.  Do this for *different groups of pixels* (corresponding to different wavelengths) to measure how the dip depth changes with wavelength.\n\n### Step 5: A Better Light Curve - \"Differential\" Photometry\n\nThe \"up-the-ramp\" effect described in the competition details is a key feature. Each observation cycle consists of a read at the beginning, an exposure time, and a read at the end. The true signal is the *difference* between the end read and the start read.\n\nLet's try to plot this differential signal. The data is provided as a continuous stream of frames. The cycle is: reset, read (start), expose, read (end), reset, read (start), ... and so on. This means the frames should come in pairs: (end, start), (end, start)...\n\n","metadata":{}},{"cell_type":"code","source":"fgs1_diff = fgs1_light_curve.iloc[1::2].values - fgs1_light_curve.iloc[0::2].values\nairs_diff = airs_light_curve.iloc[1::2].values - airs_light_curve.iloc[0::2].values\n\n\nprint(\"Plotting 'Differential' Light Curves...\")\n\n\nfig, axes = plt.subplots(nrows=2, ncols=1, figsize=(14, 9))\n\n# --- FGS1 Plot (Top subplot) ---\nsns.lineplot(\n    x=np.arange(len(fgs1_diff)), # Create an x-axis for the measurement number\n    y=fgs1_diff,\n    ax=axes[0],\n    color=\"#0077b6\", # Consistent blue color\n    lw=1 # Use a thin line for clarity\n)\naxes[0].set_title('FGS1 Differential Light Curve', fontsize=16)\naxes[0].set_ylabel('Differential Flux', fontsize=12)\n\n\n# --- AIRS-CH0 Plot (Bottom subplot) ---\nsns.lineplot(\n    x=np.arange(len(airs_diff)),\n    y=airs_diff,\n    ax=axes[1],\n    color=\"#d62728\", # Consistent red color\n    lw=1\n)\naxes[1].set_title('AIRS-CH0 Differential Light Curve', fontsize=16)\naxes[1].set_ylabel('Differential Flux', fontsize=12)\naxes[1].set_xlabel('Measurement Number (Time)', fontsize=12) # Label the shared x-axis\n\n\n\nfig.suptitle(f\"Differential Light Curves for Planet ID: {planet_id_sample}\", fontsize=20, y=1.01)\n\n# Improve layout to prevent titles/labels from overlapping\nplt.tight_layout(rect=[0, 0, 1, 0.97])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:30:22.379289Z","iopub.execute_input":"2025-06-27T20:30:22.380456Z","iopub.status.idle":"2025-06-27T20:30:23.419185Z","shell.execute_reply.started":"2025-06-27T20:30:22.380421Z","shell.execute_reply":"2025-06-27T20:30:23.418221Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n*   **AIRS-CH0 (Bottom Plot):** The transit is now stunningly clear.\n    *   **Out-of-Transit:** The flat sections before and after the dip are when the telescope is just observing the star by itself.\n    *   **Ingress/Egress:** The sharp downward and upward slopes are the beginning (ingress) and end (egress) of the planet passing in front of the star.\n    *   **In-Transit:** The bottom of the \"U\" is the period when the entire planet is in front of the star, blocking a constant amount of light.\n    *   **Noise & Artifacts:** Notice the many sharp, positive spikes. These are almost certainly cosmic ray hits on the detector. They are a secondary noise source we must remove.\n\n*   **FGS1 (Top Plot):** The signal here is more subtle but just as important.\n    *   **The Transit:** If you look closely at the cropped images, you can see a very shallow, wide dip. The FGS1 instrument is capturing the transit in visible light, and the transit depth here is different from the infrared one in AIRS-CH0.\n    *   **Systematic Trend:** The entire FGS1 light curve has a gentle, long-term curve. It's not perfectly flat outside of the transit. This is a \"systematic effect\" or \"drift,\" likely due to tiny temperature or pointing changes in the instrument over the long observation. This drift must also be modeled and removed.\n    *   **Artifacts:** The FGS1 data also has a couple of very large cosmic ray spikes.\n\n### Our Path Forward: The Data Cleaning Pipeline\n\nWe have successfully performed the first step: **differential reading**. Now we can design a pipeline to address the remaining issues before we can measure the transit depth.\n\nOur next goal is to create a **clean, normalized light curve**. The steps are:\n1.  **Artifact Removal (Sigma Clipping):** Remove the sharp cosmic ray spikes.\n2.  **Detrending & Normalization:** Remove the long-term systematic drift (like the curve in FGS1) and scale the light curve so the out-of-transit flux is exactly 1.0.\n\nLet's do this for the AIRS-CH0 data, as its features are most obvious.\n\n### Step 6: Cleaning and Normalizing a Light Curve\n\n","metadata":{}},{"cell_type":"code","source":"from scipy.signal import medfilt\n\n\ndef sigma_clip(data, window_size=51, sigma=5):\n    \"\"\"Simple sigma clipping function.\"\"\"\n    local_median = medfilt(data, kernel_size=window_size)\n    residual = data - local_median\n    mad = np.median(np.abs(residual))\n    robust_std = mad * 1.4826\n    outliers = np.abs(residual) > (sigma * robust_std)\n    clean_data = data.copy()\n    clean_data[outliers] = np.nan\n    return clean_data, outliers\n\nprint(\"--- 1. Performing Sigma Clipping on AIRS-CH0 Data ---\")\nairs_clipped, airs_outliers = sigma_clip(airs_diff)\n\n\n# --- 2. Detrending and Normalization - This code is unchanged ---\nout_of_transit_mask = np.ones_like(airs_clipped, dtype=bool)\nout_of_transit_mask[1700:4200] = False\nout_of_transit_mask[np.isnan(airs_clipped)] = False\nx_axis = np.arange(len(airs_clipped))\npoly_coeffs = np.polyfit(x_axis[out_of_transit_mask], airs_clipped[out_of_transit_mask], 2)\nbaseline_model = np.polyval(poly_coeffs, x_axis)\nairs_normalized = airs_clipped / baseline_model\n\n\n# --- 3. Visualize the Entire Cleaning Process (Seaborn/Matplotlib version) ---\nprint(\"\\n--- 2. Visualizing the Cleaning Pipeline ---\")\n\n# Create a figure with 2 subplots, one on top of the other, sharing the x-axis\nfig, axes = plt.subplots(nrows=2, ncols=1, figsize=(14, 10), sharex=True)\n\n# --- Top Plot: Sigma Clipping and Baseline Fitting ---\nax_top = axes[0]\n# Plot 1: The original differential data as a light background line\nsns.lineplot(x=x_axis, y=airs_diff, ax=ax_top, color='lightgrey', label='Original Diff. Data')\n# Plot 2: The baseline model as a thick orange line\nsns.lineplot(x=x_axis, y=baseline_model, ax=ax_top, color='orange', linewidth=3, label='Baseline Fit')\n# Plot 3: The clipped outliers as red markers\n# Note: we plot the original 'airs_diff' values at the outlier indices\nsns.scatterplot(x=x_axis[airs_outliers], y=airs_diff[airs_outliers], ax=ax_top, color='red', marker='o', s=50, label='Clipped Outliers', zorder=5)\n\nax_top.set_title('Step 1: Sigma Clipping and Baseline Fitting', fontsize=16)\nax_top.set_ylabel('Flux', fontsize=12)\nax_top.legend()\n\n\n# --- Bottom Plot: Final Normalized Light Curve ---\nax_bottom = axes[1]\nsns.lineplot(x=x_axis, y=airs_normalized, ax=ax_bottom, color='dodgerblue', label='Normalized Flux')\nax_bottom.set_title('Step 2: Final Normalized Light Curve', fontsize=16)\nax_bottom.set_ylabel('Normalized Flux', fontsize=12)\nax_bottom.set_xlabel('Measurement Number', fontsize=12)\n# Set y-axis limits to better see the transit depth\nax_bottom.set_ylim(bottom=min(0.98, np.nanmin(airs_normalized) - 0.005), top=1.01)\nax_bottom.legend()\n\n\nfig.suptitle(f\"Data Cleaning Pipeline for Planet {planet_id_sample} (AIRS-CH0)\", fontsize=20, y=1.0)\nplt.tight_layout(rect=[0, 0, 1, 0.97])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:31:47.486895Z","iopub.execute_input":"2025-06-27T20:31:47.487291Z","iopub.status.idle":"2025-06-27T20:31:48.326588Z","shell.execute_reply.started":"2025-06-27T20:31:47.487263Z","shell.execute_reply":"2025-06-27T20:31:48.325614Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n### Analysis of the Cleaning Pipeline\n\n*   **Step 1 (Top Plot):** Our sigma clipping (red dots) has perfectly identified and ignored the cosmic ray artifacts. The baseline fit (orange line) has correctly identified the stable, out-of-transit flux level. This process is robust.\n*   **Step 2 (Bottom Plot):** This is the result of our hard work. A beautiful, normalized light curve. The y-axis is now directly interpretable: the flux is 1.0 when the star is unobscured. The depth of the transit is now trivial to measure: it's simply `1.0 - (the flux at the bottom of the dip)`. In this case, the transit depth is roughly `1.0 - 0.982 = 0.018`.\n\nWe have successfully built a prototype data reduction pipeline.\n\n### The Grand Strategy: From Clean Light Curves to a Winning Model\n\nSo, how do we get from this one clean light curve to a competition submission? The key is realizing that the ground truth spectrum we saw in Step 3 is just a collection of these transit depths measured at different wavelengths.\n\nHere is the complete end-to-end strategy:\n\n**Part 1: The Data Processing Pipeline (Scaling Up What We Just Did)**\n\nThe core idea is to generate a \"measured spectrum\" for every planet in the dataset by running our cleaning pipeline on each wavelength channel individually.\n\n1.  **For each planet:**\n2.  **For each of the 283 wavelength channels:**\n    a.  **Select Pixels:** Isolate the raw signal data for only the pixels corresponding to that specific wavelength.\n    b.  **Create Light Curve:** Generate a light curve for that channel (using differential reads).\n    c.  **Clean & Normalize:** Run the full cleaning pipeline (sigma clip, detrend, normalize) on that one channel's light curve.\n    d.  **Measure Transit Depth:** From the final clean light curve, calculate the mean flux during the transit. The depth is `1.0 - mean_in_transit_flux`. This value is your *measured* `μ` for that channel.\n    e.  **Estimate Uncertainty:** Calculate the standard deviation of the flux during the transit. This can be your initial estimate for `σ` for that channel.\n3.  **Result:** After looping through all 283 channels, you will have assembled one \"measured spectrum\" and its corresponding uncertainties for that planet.\n\n**Part 2: The Machine Learning Model (Denoising and Predicting)**\n\nThe \"measured spectrum\" from Part 1 will still be noisy. The goal of the ML model is to use the physical properties of the system to predict a cleaner, more accurate spectrum.\n\n*   **Features (X):** The star and planet physical parameters from `train_star_info.csv` (`Rs`, `Ms`, `Ts`, `Mp`, etc.).\n*   **Target (y):** The ground truth spectrum from `train.csv`.\n\nYou would train a model that learns the mapping: `f(physical_parameters) -> true_spectrum`. This could be:\n*   A Multi-Layer Perceptron (MLP) with ~8 input neurons and 283 output neurons.\n*   A Gradient Boosting model like LightGBM or XGBoost, trained to predict each of the 283 wavelength channels independently.\n*   More advanced architectures like Transformers or CNNs that can treat the spectrum as a sequence.\n\n**Part 3: The Final Prediction**\n\nFor the hidden test set, your notebook will:\n1.  Receive the raw data for a test planet.\n2.  Run your **Part 1 Pipeline** on it to generate a noisy \"measured spectrum\".\n3.  Feed the planet's physical parameters (from `test_star_info.csv`) into your trained **Part 2 ML Model** to get a clean, predicted spectrum (`μ_user`).\n4.  Estimate the uncertainty (`σ_user`). This is a critical and difficult step. It could come from the uncertainty of your ML model's prediction, combined with the measurement uncertainty from your pipeline.\n5.  Write the `μ_user` and `σ_user` values to `submission.csv`.\n","metadata":{}},{"cell_type":"markdown","source":"# Next part upcoming...","metadata":{}},{"cell_type":"code","source":"print(\"--- Examining the Raw Data Structure ---\")\nprint(f\"Shape of FGS1 signal data: {fgs1_raw_df.shape}\")\nprint(f\"Expected pixel count for a 32x32 detector: 32 * 32 = {32*32}\\n\")\n\nprint(f\"Shape of AIRS-CH0 signal data: {airs_raw_df.shape}\")\nprint(f\"Expected pixel count for a 32x356 detector: 32 * 356 = {32*356}\\n\")\n\ngain = adc_info_df['FGS1_adc_gain'].iloc[0]\noffset = adc_info_df['FGS1_adc_offset'].iloc[0]\n\n# FGS1 Data\nframe_index_fgs1 = 5000\nfgs1_frame_flat = fgs1_raw_df.iloc[frame_index_fgs1]\nfgs1_frame_corrected = fgs1_frame_flat * gain + offset\nfgs1_frame_2d = fgs1_frame_corrected.values.reshape(32, 32)\n\n# AIRS-CH0 Data\nframe_index_airs = 5000\nairs_frame_flat = airs_raw_df.iloc[frame_index_airs]\nairs_frame_corrected = airs_frame_flat * gain + offset\nairs_frame_2d = airs_frame_corrected.values.reshape(32, 356)\n\n\n\nprint(\"--- Visualizing Detector Frames with Seaborn ---\")\n\n\n# We make the figure wider to accommodate the long AIRS-CH0 detector\nfig, axes = plt.subplots(nrows=2, ncols=1, figsize=(15, 10))\n\n# --- FGS1 Visualization (Top Plot) ---\nax_fgs1 = axes[0]\nsns.heatmap(\n    fgs1_frame_2d,\n    ax=ax_fgs1,\n    cmap='viridis', # This is the same color scale as in Plotly\n    cbar_kws={'label': 'Flux'} # Add a label to the color bar\n)\nax_fgs1.set_title(f'FGS1 Detector Frame #{frame_index_fgs1}', fontsize=16)\nax_fgs1.set_xlabel('Pixel Column', fontsize=12)\nax_fgs1.set_ylabel('Pixel Row', fontsize=12)\n\n# --- AIRS-CH0 Visualization (Bottom Plot) ---\nax_airs = axes[1]\nsns.heatmap(\n    airs_frame_2d,\n    ax=ax_airs,\n    cmap='viridis',\n    cbar_kws={'label': 'Flux'}\n)\nax_airs.set_title(f'AIRS-CH0 Detector Frame #{frame_index_airs}', fontsize=16)\nax_airs.set_xlabel('Wavelength Axis (Pixel Column)', fontsize=12)\nax_airs.set_ylabel('Spatial Axis (Pixel Row)', fontsize=12)\n\n\n# --- Final Touches ---\nfig.suptitle(f\"Single Detector Frames for Planet {planet_id_sample}\", fontsize=20, y=1.0)\nplt.tight_layout(rect=[0, 0, 1, 0.97])\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T20:32:54.629138Z","iopub.execute_input":"2025-06-27T20:32:54.629589Z","iopub.status.idle":"2025-06-27T20:32:56.86363Z","shell.execute_reply.started":"2025-06-27T20:32:54.62956Z","shell.execute_reply":"2025-06-27T20:32:56.86263Z"},"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null}]}