{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\n# Use the kagglehub client library to attach Kaggle resources like competitions, datasets, and models to your session\n# Learn more about kagglehub: https://github.com/Kaggle/kagglehub/blob/main/README.md\n\nimport kagglehub\n# kagglehub.dataset_download('<owner>/<dataset-slug>')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:29:45.276745Z","iopub.execute_input":"2026-09-21T02:29:45.277211Z","iopub.status.idle":"2026-09-21T02:30:24.880334Z","shell.execute_reply.started":"2026-09-21T02:29:45.277173Z","shell.execute_reply":"2026-09-21T02:30:24.878556Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **1) Research Objective & Notebook Introduction**","metadata":{}},{"cell_type":"markdown","source":"## **Ariel Data Challenge 2025**\n## Exploratory Data Analysis & Data Audit\n\n### Research Objective\n\nThe objective of this notebook is to perform a systematic exploratory\nanalysis of the Ariel Data Challenge dataset before applying detector\ncalibration, spectral extraction, denoising, and machine-learning methods.\n\nThe analysis focuses on understanding:\n\n- The structure and hierarchy of the Ariel dataset\n- Planet-level and visit-level observations\n- Detector signal dimensions and characteristics\n- Stellar and observational metadata\n- Missing values and invalid measurements\n- Signal distributions and outliers\n- Temporal and spectral behaviour\n- Detector-level spatial patterns\n- Differences between training and test data\n- Initial estimates of signal-to-noise characteristics\n\nThe primary goal is to identify the statistical and instrumental\nproperties of the data that should influence the subsequent preprocessing\npipeline.\n\n---\n\n### Notebook Scope\n\nThis notebook intentionally focuses on **EDA and data auditing**.\n\nThe following operations are reserved for subsequent preprocessing\nnotebooks:\n\n1. ADC calibration\n2. Linearity correction\n3. Dark-current subtraction\n4. Correlated double sampling (CDS)\n5. Flat-field correction\n6. Dead/hot pixel handling\n7. Spectral trace extraction\n8. Jitter correction\n9. Temporal and spectral binning\n10. Transit modelling\n11. Transmission-spectrum estimation\n\nThe output of this notebook will be a set of documented observations,\ndata-quality findings, and preprocessing hypotheses that will guide\nthose subsequent stages.\n\n---\n\n### Scientific Pipeline\n\nRaw Ariel Data\n       │\n       ▼\nData Audit & EDA              ← This notebook\n       │\n       ▼\nDetector Calibration\n       │\n       ▼\nSpectral Extraction\n       │\n       ▼\nDenoising & Systematics Correction\n       │\n       ▼\nTransit Characterization\n       │\n       ▼\nTransmission Spectrum\n       │\n       ▼\nMachine Learning / Uncertainty Calibration\n       │\n       ▼\nFinal Prediction","metadata":{}},{"cell_type":"markdown","source":"**Research Questions**\n\nThis notebook will investigate the following questions:\n\nRQ1. What is the hierarchical structure of the Ariel dataset?\n\nRQ2. What are the dimensions and numerical characteristics of the\nraw detector signals?\n\nRQ3. Are there missing, invalid, saturated, or anomalous measurements?\n\nRQ4. How do signal characteristics vary across planets and visits?\n\nRQ5. What spatial structures are present in the detector data?\n\nRQ6. What temporal patterns and systematic variations are present?\n\nRQ7. How are the stellar and observational parameters distributed?\n\nRQ8. How similar are the training and test populations?\n\nRQ9. What evidence exists for white versus correlated noise?\n\nRQ10. Which properties of the dataset should influence the design of\nthe preprocessing and validation pipeline?","metadata":{}},{"cell_type":"markdown","source":"# **2) Environment & Reproducibility**","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport json\nimport random\nimport warnings\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom IPython.display import display\n\nwarnings.filterwarnings(\"ignore\")\n\n\n# Reproducibility\n\n\nSEED = 42\n\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\n\n\n\n\npd.set_option(\"display.max_columns\", 100)\npd.set_option(\"display.max_rows\", 100)\npd.set_option(\"display.float_format\", lambda x: f\"{x:.6g}\")\n\n\n# Plot configuration\n\n\nplt.rcParams[\"figure.figsize\"] = (10, 6)\nplt.rcParams[\"axes.grid\"] = True\n\nprint(\"Environment initialized.\")\nprint(f\"Random seed: {SEED}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:30:47.429324Z","iopub.execute_input":"2026-09-21T02:30:47.42973Z","iopub.status.idle":"2026-09-21T02:30:47.438661Z","shell.execute_reply.started":"2026-09-21T02:30:47.429694Z","shell.execute_reply":"2026-09-21T02:30:47.437549Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n# KAGGLE ENVIRONMENT\n\n\nimport platform\nimport sys\n\nprint(\"Python version :\", sys.version)\nprint(\"Platform       :\", platform.platform())\nprint(\"NumPy version  :\", np.__version__)\nprint(\"Pandas version :\", pd.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:30:52.335384Z","iopub.execute_input":"2026-09-21T02:30:52.335831Z","iopub.status.idle":"2026-09-21T02:30:52.342642Z","shell.execute_reply.started":"2026-09-21T02:30:52.3358Z","shell.execute_reply":"2026-09-21T02:30:52.341731Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import shutil\n\ntotal, used, free = shutil.disk_usage(\"/kaggle/working\")\n\nprint(f\"Disk total : {total / 1024**3:.2f} GB\")\nprint(f\"Disk used  : {used / 1024**3:.2f} GB\")\nprint(f\"Disk free  : {free / 1024**3:.2f} GB\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:30:55.322311Z","iopub.execute_input":"2026-09-21T02:30:55.323228Z","iopub.status.idle":"2026-09-21T02:30:55.331002Z","shell.execute_reply.started":"2026-09-21T02:30:55.323186Z","shell.execute_reply":"2026-09-21T02:30:55.329861Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **3) Dataset Discovery**","metadata":{}},{"cell_type":"code","source":"KAGGLE_INPUT = Path(\"/kaggle/input\")\nKAGGLE_WORKING = Path(\"/kaggle/working\")\n\nprint(\"Kaggle input directories:\")\nfor path in sorted(KAGGLE_INPUT.iterdir()):\n    print(\" -\", path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:32:39.466468Z","iopub.execute_input":"2026-09-21T02:32:39.467793Z","iopub.status.idle":"2026-09-21T02:32:39.475081Z","shell.execute_reply.started":"2026-09-21T02:32:39.467738Z","shell.execute_reply":"2026-09-21T02:32:39.473607Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\nfiles = []\n\nfor path in KAGGLE_INPUT.rglob(\"*\"):\n    if path.is_file():\n        files.append({\n            \"path\": str(path),\n            \"name\": path.name,\n            \"extension\": path.suffix.lower(),\n            \"size_mb\": path.stat().st_size / 1024**2\n        })\n\nfile_inventory = pd.DataFrame(files)\n\nprint(f\"Total files discovered: {len(file_inventory)}\")\n\ndisplay(\n    file_inventory\n    .sort_values(\"size_mb\", ascending=False)\n    .head(50)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:31:01.482416Z","iopub.execute_input":"2026-09-21T02:31:01.483151Z","iopub.status.idle":"2026-09-21T02:31:54.872256Z","shell.execute_reply.started":"2026-09-21T02:31:01.483106Z","shell.execute_reply":"2026-09-21T02:31:54.870704Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\nextension_summary = (\n    file_inventory\n    .groupby(\"extension\")\n    .agg(\n        file_count=(\"name\", \"count\"),\n        total_size_mb=(\"size_mb\", \"sum\")\n    )\n    .sort_values(\"total_size_mb\", ascending=False)\n)\n\ndisplay(extension_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:32:18.041813Z","iopub.execute_input":"2026-09-21T02:32:18.043429Z","iopub.status.idle":"2026-09-21T02:32:18.079429Z","shell.execute_reply.started":"2026-09-21T02:32:18.043381Z","shell.execute_reply":"2026-09-21T02:32:18.077721Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# SEARCH FOR IMPORTANT ARIEL FILES / DIRECTORIES\n# ============================================================\n\nkeywords = [\n    \"adc\",\n    \"axis\",\n    \"star\",\n    \"train\",\n    \"test\",\n    \"signal\",\n    \"label\"\n]\n\nfor keyword in keywords:\n    matches = [\n        path for path in file_inventory[\"path\"]\n        if keyword.lower() in path.lower()\n    ]\n\n    print(f\"\\n[{keyword.upper()}]\")\n    \n    for match in matches[:20]:\n        print(match)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:32:45.510601Z","iopub.execute_input":"2026-09-21T02:32:45.511808Z","iopub.status.idle":"2026-09-21T02:32:45.562671Z","shell.execute_reply.started":"2026-09-21T02:32:45.511762Z","shell.execute_reply":"2026-09-21T02:32:45.561513Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **4) Dataset Manifest & Structural Audit**","metadata":{}},{"cell_type":"markdown","source":"## **4.1) Set Root**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 5. DATASET MANIFEST & STRUCTURAL AUDIT\n# ============================================================\n\nfrom pathlib import Path\n\nARIEL_ROOT = Path(\n    \"/kaggle/input/competitions/ariel-data-challenge-2025\"\n)\n\nTRAIN_ROOT = ARIEL_ROOT / \"train\"\nTEST_ROOT = ARIEL_ROOT / \"test\"\n\nprint(\"Ariel dataset:\")\nprint(ARIEL_ROOT)\n\nprint(\"\\nTrain directory:\")\nprint(TRAIN_ROOT)\n\nprint(\"\\nTest directory:\")\nprint(TEST_ROOT)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:32:49.985552Z","iopub.execute_input":"2026-09-21T02:32:49.986183Z","iopub.status.idle":"2026-09-21T02:32:49.99421Z","shell.execute_reply.started":"2026-09-21T02:32:49.986142Z","shell.execute_reply":"2026-09-21T02:32:49.992758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.2) Identify Planets**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# PLANET IDS\n# ============================================================\n\ntrain_planets = sorted([\n    p.name for p in TRAIN_ROOT.iterdir()\n    if p.is_dir()\n])\n\ntest_planets = sorted([\n    p.name for p in TEST_ROOT.iterdir()\n    if p.is_dir()\n])\n\nprint(f\"Number of training planets : {len(train_planets):,}\")\nprint(f\"Number of test planets     : {len(test_planets):,}\")\n\nprint(\"\\nFirst 10 training planets:\")\nprint(train_planets[:10])\n\nprint(\"\\nFirst 10 test planets:\")\nprint(test_planets[:10])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:32:54.666204Z","iopub.execute_input":"2026-09-21T02:32:54.666611Z","iopub.status.idle":"2026-09-21T02:32:55.711382Z","shell.execute_reply.started":"2026-09-21T02:32:54.666582Z","shell.execute_reply":"2026-09-21T02:32:55.710454Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.3) Planet Level Manifest**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# PLANET-LEVEL MANIFEST\n# ============================================================\n\ndef build_planet_manifest(root, split):\n    records = []\n\n    for planet_dir in sorted(root.iterdir()):\n        if not planet_dir.is_dir():\n            continue\n\n        planet_id = planet_dir.name\n\n        files = list(planet_dir.rglob(\"*\"))\n\n        signal_files = [\n            f for f in files\n            if f.is_file() and \"signal\" in f.name\n        ]\n\n        calibration_files = [\n            f for f in files\n            if f.is_file() and \"calibration\" in str(f)\n        ]\n\n        records.append({\n            \"split\": split,\n            \"planet_id\": planet_id,\n            \"n_signal_files\": len(signal_files),\n            \"n_calibration_files\": len(calibration_files),\n            \"n_total_files\": len([\n                f for f in files if f.is_file()\n            ])\n        })\n\n    return pd.DataFrame(records)\n\n\ntrain_manifest = build_planet_manifest(\n    TRAIN_ROOT,\n    \"train\"\n)\n\ntest_manifest = build_planet_manifest(\n    TEST_ROOT,\n    \"test\"\n)\n\nmanifest = pd.concat(\n    [train_manifest, test_manifest],\n    ignore_index=True\n)\n\ndisplay(manifest.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:33:09.697821Z","iopub.execute_input":"2026-09-21T02:33:09.698883Z","iopub.status.idle":"2026-09-21T02:33:33.008216Z","shell.execute_reply.started":"2026-09-21T02:33:09.698792Z","shell.execute_reply":"2026-09-21T02:33:33.006955Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.4) Manifest Summary**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# MANIFEST SUMMARY\n# ============================================================\n\nprint(\"Manifest shape:\", manifest.shape)\n\ndisplay(\n    manifest.groupby(\"split\").agg(\n        planets=(\"planet_id\", \"nunique\"),\n        total_signal_files=(\"n_signal_files\", \"sum\"),\n        total_calibration_files=(\"n_calibration_files\", \"sum\"),\n        total_files=(\"n_total_files\", \"sum\")\n    )\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:33:38.798489Z","iopub.execute_input":"2026-09-21T02:33:38.799225Z","iopub.status.idle":"2026-09-21T02:33:38.830164Z","shell.execute_reply.started":"2026-09-21T02:33:38.799175Z","shell.execute_reply":"2026-09-21T02:33:38.828675Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.5) File Dsitribution per Planet**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# FILE DISTRIBUTION PER PLANET\n# ============================================================\n\nfig, axes = plt.subplots(1, 3, figsize=(18, 5))\n\nfor split, group in manifest.groupby(\"split\"):\n\n    axes[0].hist(\n        group[\"n_signal_files\"],\n        bins=20,\n        alpha=0.6,\n        label=split\n    )\n\n    axes[1].hist(\n        group[\"n_calibration_files\"],\n        bins=20,\n        alpha=0.6,\n        label=split\n    )\n\n    axes[2].hist(\n        group[\"n_total_files\"],\n        bins=20,\n        alpha=0.6,\n        label=split\n    )\n\naxes[0].set_title(\"Signal files per planet\")\naxes[1].set_title(\"Calibration files per planet\")\naxes[2].set_title(\"Total files per planet\")\n\nfor ax in axes:\n    ax.set_xlabel(\"Number of files\")\n    ax.set_ylabel(\"Number of planets\")\n    ax.legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:33:49.151838Z","iopub.execute_input":"2026-09-21T02:33:49.1531Z","iopub.status.idle":"2026-09-21T02:33:50.00009Z","shell.execute_reply.started":"2026-09-21T02:33:49.153043Z","shell.execute_reply":"2026-09-21T02:33:49.999015Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.6)Single Planet Structure**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SINGLE PLANET STRUCTURE\n# ============================================================\n\nsample_planet = train_planets[0]\n\nsample_dir = TRAIN_ROOT / sample_planet\n\nprint(f\"Planet ID: {sample_planet}\")\nprint(\"=\" * 70)\n\nfor path in sorted(sample_dir.rglob(\"*\")):\n    if path.is_file():\n        print(path.relative_to(sample_dir))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:00.369441Z","iopub.execute_input":"2026-09-21T02:34:00.370595Z","iopub.status.idle":"2026-09-21T02:34:00.400034Z","shell.execute_reply.started":"2026-09-21T02:34:00.37054Z","shell.execute_reply":"2026-09-21T02:34:00.39839Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.7)Signal File Types**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SIGNAL FILE TYPES\n# ============================================================\n\nsignal_file_records = []\n\nfor split, root in [\n    (\"train\", TRAIN_ROOT),\n    (\"test\", TEST_ROOT)\n]:\n    \n    for planet_dir in root.iterdir():\n\n        if not planet_dir.is_dir():\n            continue\n\n        for file in planet_dir.rglob(\"*_signal_*.parquet\"):\n\n            name = file.name\n\n            if name.startswith(\"AIRS\"):\n                instrument = \"AIRS\"\n            elif name.startswith(\"FGS1\"):\n                instrument = \"FGS1\"\n            else:\n                instrument = \"Other\"\n\n            signal_file_records.append({\n                \"split\": split,\n                \"planet_id\": planet_dir.name,\n                \"file\": name,\n                \"instrument\": instrument\n            })\n\n\nsignal_manifest = pd.DataFrame(signal_file_records)\n\ndisplay(\n    signal_manifest\n    .groupby([\"split\", \"instrument\"])\n    .size()\n    .rename(\"file_count\")\n    .reset_index()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:04.564981Z","iopub.execute_input":"2026-09-21T02:34:04.565495Z","iopub.status.idle":"2026-09-21T02:34:12.910389Z","shell.execute_reply.started":"2026-09-21T02:34:04.565454Z","shell.execute_reply":"2026-09-21T02:34:12.909242Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.8) Signal File Distribution**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SIGNAL FILE DISTRIBUTION\n# ============================================================\n\nsignal_summary = (\n    signal_manifest\n    .groupby([\"split\", \"instrument\"])\n    .size()\n    .unstack(fill_value=0)\n)\n\ndisplay(signal_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:16.057273Z","iopub.execute_input":"2026-09-21T02:34:16.058265Z","iopub.status.idle":"2026-09-21T02:34:16.088728Z","shell.execute_reply.started":"2026-09-21T02:34:16.058221Z","shell.execute_reply":"2026-09-21T02:34:16.087764Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.9) Identification of Visit**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# VISIT IDENTIFICATION\n# ============================================================\n\nsignal_manifest[\"visit\"] = (\n    signal_manifest[\"file\"]\n    .str.extract(r\"_signal_(\\d+)\\.parquet\")[0]\n    .astype(\"Int64\")\n)\n\ndisplay(signal_manifest.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:23.523719Z","iopub.execute_input":"2026-09-21T02:34:23.524947Z","iopub.status.idle":"2026-09-21T02:34:23.548111Z","shell.execute_reply.started":"2026-09-21T02:34:23.524888Z","shell.execute_reply":"2026-09-21T02:34:23.546683Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# VISITS PER PLANET\n# ============================================================\n\nvisit_summary = (\n    signal_manifest\n    .groupby([\"split\", \"planet_id\"])[\"visit\"]\n    .nunique()\n    .reset_index(name=\"n_visits\")\n)\n\ndisplay(visit_summary.head(20))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:33.045999Z","iopub.execute_input":"2026-09-21T02:34:33.047155Z","iopub.status.idle":"2026-09-21T02:34:33.069122Z","shell.execute_reply.started":"2026-09-21T02:34:33.047107Z","shell.execute_reply":"2026-09-21T02:34:33.067918Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# VISIT COUNT DISTRIBUTION\n# ============================================================\n\nplt.figure(figsize=(10, 6))\n\nfor split, group in visit_summary.groupby(\"split\"):\n    plt.hist(\n        group[\"n_visits\"],\n        bins=np.arange(\n            group[\"n_visits\"].min(),\n            group[\"n_visits\"].max() + 2\n        ) - 0.5,\n        alpha=0.6,\n        label=split\n    )\n\nplt.xlabel(\"Number of visits\")\nplt.ylabel(\"Number of planets\")\nplt.title(\"Distribution of Visits per Planet\")\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:37.040247Z","iopub.execute_input":"2026-09-21T02:34:37.041034Z","iopub.status.idle":"2026-09-21T02:34:37.293242Z","shell.execute_reply.started":"2026-09-21T02:34:37.040984Z","shell.execute_reply":"2026-09-21T02:34:37.291908Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.10) Metadata**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# LOAD METADATA\n# ============================================================\n\ntrain_labels = pd.read_csv(\n    ARIEL_ROOT / \"train.csv\"\n)\n\ntrain_star_info = pd.read_csv(\n    ARIEL_ROOT / \"train_star_info.csv\"\n)\n\ntest_star_info = pd.read_csv(\n    ARIEL_ROOT / \"test_star_info.csv\"\n)\n\nadc_info = pd.read_csv(\n    ARIEL_ROOT / \"adc_info.csv\"\n)\n\naxis_info = pd.read_parquet(\n    ARIEL_ROOT / \"axis_info.parquet\"\n)\n\nprint(\"train.csv:\")\nprint(train_labels.shape)\n\nprint(\"\\ntrain_star_info.csv:\")\nprint(train_star_info.shape)\n\nprint(\"\\ntest_star_info.csv:\")\nprint(test_star_info.shape)\n\nprint(\"\\nadc_info.csv:\")\nprint(adc_info.shape)\n\nprint(\"\\naxis_info.parquet:\")\nprint(axis_info.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:45.612129Z","iopub.execute_input":"2026-09-21T02:34:45.612515Z","iopub.status.idle":"2026-09-21T02:34:46.225502Z","shell.execute_reply.started":"2026-09-21T02:34:45.612479Z","shell.execute_reply":"2026-09-21T02:34:46.223695Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.11)Column Inspection**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# METADATA COLUMN AUDIT\n# ============================================================\n\nmetadata_tables = {\n    \"train_labels\": train_labels,\n    \"train_star_info\": train_star_info,\n    \"test_star_info\": test_star_info,\n    \"adc_info\": adc_info,\n    \"axis_info\": axis_info\n}\n\nfor name, df in metadata_tables.items():\n\n    print(\"\\n\" + \"=\" * 80)\n    print(name)\n    print(\"=\" * 80)\n\n    print(\"Shape:\", df.shape)\n    print(\"Columns:\")\n\n    for col in df.columns:\n        print(\"  -\", col)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:34:50.489463Z","iopub.execute_input":"2026-09-21T02:34:50.490035Z","iopub.status.idle":"2026-09-21T02:34:50.501766Z","shell.execute_reply.started":"2026-09-21T02:34:50.489991Z","shell.execute_reply":"2026-09-21T02:34:50.50022Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.12) Preview**","metadata":{}},{"cell_type":"code","source":"for name, df in metadata_tables.items():\n\n    print(\"\\n\" + \"=\" * 80)\n    print(name)\n    print(\"=\" * 80)\n\n    display(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:04.631577Z","iopub.execute_input":"2026-09-21T02:35:04.632321Z","iopub.status.idle":"2026-09-21T02:35:04.742316Z","shell.execute_reply.started":"2026-09-21T02:35:04.632275Z","shell.execute_reply":"2026-09-21T02:35:04.741199Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.13) Memory Usage**","metadata":{}},{"cell_type":"markdown","source":"## ","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# METADATA MEMORY USAGE\n# ============================================================\n\nmemory_summary = []\n\nfor name, df in metadata_tables.items():\n\n    memory_mb = (\n        df.memory_usage(deep=True).sum()\n        / 1024**2\n    )\n\n    memory_summary.append({\n        \"dataset\": name,\n        \"rows\": len(df),\n        \"columns\": len(df.columns),\n        \"memory_mb\": memory_mb\n    })\n\nmemory_summary = pd.DataFrame(memory_summary)\n\ndisplay(memory_summary)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:13.019452Z","iopub.execute_input":"2026-09-21T02:35:13.020157Z","iopub.status.idle":"2026-09-21T02:35:13.224357Z","shell.execute_reply.started":"2026-09-21T02:35:13.020108Z","shell.execute_reply":"2026-09-21T02:35:13.223006Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.14) Master Planet Manifest**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# MASTER PLANET MANIFEST\n# ============================================================\n\nplanet_manifest = (\n    manifest\n    .merge(\n        visit_summary,\n        on=[\"split\", \"planet_id\"],\n        how=\"left\"\n    )\n)\n\nplanet_manifest[\"planet_id\"] = (\n    planet_manifest[\"planet_id\"]\n    .astype(str)\n)\n\nprint(\"Master manifest shape:\", planet_manifest.shape)\n\ndisplay(planet_manifest.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:37.733274Z","iopub.execute_input":"2026-09-21T02:35:37.734347Z","iopub.status.idle":"2026-09-21T02:35:37.764909Z","shell.execute_reply.started":"2026-09-21T02:35:37.734303Z","shell.execute_reply":"2026-09-21T02:35:37.763672Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **4.15) Filter Lightweight Manifest**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SAVE LIGHTWEIGHT MANIFEST\n# ============================================================\n\nMANIFEST_PATH = (\n    KAGGLE_WORKING / \"ariel_planet_manifest.parquet\"\n)\n\nplanet_manifest.to_parquet(\n    MANIFEST_PATH,\n    index=False\n)\n\nprint(f\"Saved manifest to:\\n{MANIFEST_PATH}\")\nprint(\n    f\"Size: {MANIFEST_PATH.stat().st_size / 1024**2:.2f} MB\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:41.003806Z","iopub.execute_input":"2026-09-21T02:35:41.004398Z","iopub.status.idle":"2026-09-21T02:35:41.029427Z","shell.execute_reply.started":"2026-09-21T02:35:41.004363Z","shell.execute_reply":"2026-09-21T02:35:41.028185Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **5) Metadata EDA**","metadata":{}},{"cell_type":"markdown","source":"## **5.1) Load Metadata EDA**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 6 — METADATA EDA\n# 6.1 Load metadata tables\n# ============================================================\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nfrom pathlib import Path\n\nARIEL_ROOT = Path(\"/kaggle/input/competitions/ariel-data-challenge-2025\")\n\n# ------------------------------------------------------------\n# Load small metadata tables\n# ------------------------------------------------------------\n\ntrain_labels = pd.read_csv(ARIEL_ROOT / \"train.csv\")\ntrain_star_info = pd.read_csv(ARIEL_ROOT / \"train_star_info.csv\")\ntest_star_info = pd.read_csv(ARIEL_ROOT / \"test_star_info.csv\")\nadc_info = pd.read_csv(ARIEL_ROOT / \"adc_info.csv\")\naxis_info = pd.read_parquet(ARIEL_ROOT / \"axis_info.parquet\")\n\nmetadata = {\n    \"train.csv\": train_labels,\n    \"train_star_info.csv\": train_star_info,\n    \"test_star_info.csv\": test_star_info,\n    \"adc_info.csv\": adc_info,\n    \"axis_info.parquet\": axis_info,\n}\n\nprint(\"=\" * 80)\nprint(\"METADATA TABLE OVERVIEW\")\nprint(\"=\" * 80)\n\nfor name, df in metadata.items():\n    print(f\"\\n{name}\")\n    print(\"-\" * 60)\n    print(\"Shape:\", df.shape)\n    print(\"Columns:\", list(df.columns))\n    print(\"Memory:\", f\"{df.memory_usage(deep=True).sum() / 1024**2:.3f} MB\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:47.220999Z","iopub.execute_input":"2026-09-21T02:35:47.221683Z","iopub.status.idle":"2026-09-21T02:35:47.531523Z","shell.execute_reply.started":"2026-09-21T02:35:47.221636Z","shell.execute_reply":"2026-09-21T02:35:47.529812Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.2) Table Inspection**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.2 — Basic previews\n# ============================================================\n\nfor name, df in metadata.items():\n    print(\"\\n\" + \"=\" * 100)\n    print(name)\n    print(\"=\" * 100)\n\n    display(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:35:57.681914Z","iopub.execute_input":"2026-09-21T02:35:57.683985Z","iopub.status.idle":"2026-09-21T02:35:57.772225Z","shell.execute_reply.started":"2026-09-21T02:35:57.683924Z","shell.execute_reply":"2026-09-21T02:35:57.77095Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Data types\n# ============================================================\n\nfor name, df in metadata.items():\n    print(\"\\n\" + \"=\" * 80)\n    print(name)\n    print(\"=\" * 80)\n\n    display(\n        pd.DataFrame({\n            \"dtype\": df.dtypes.astype(str),\n            \"non_null\": df.notna().sum(),\n            \"null_count\": df.isna().sum(),\n            \"unique_values\": df.nunique(dropna=True)\n        })\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:36:04.779256Z","iopub.execute_input":"2026-09-21T02:36:04.779977Z","iopub.status.idle":"2026-09-21T02:36:04.897342Z","shell.execute_reply.started":"2026-09-21T02:36:04.779932Z","shell.execute_reply":"2026-09-21T02:36:04.896129Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.3) Scan for missing values**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.3 — Missing-value analysis\n# ============================================================\n\nmissing_summary = []\n\nfor name, df in metadata.items():\n\n    total_cells = df.shape[0] * df.shape[1]\n\n    for col in df.columns:\n        missing_count = df[col].isna().sum()\n\n        missing_summary.append({\n            \"table\": name,\n            \"column\": col,\n            \"missing_count\": missing_count,\n            \"missing_fraction\": missing_count / len(df),\n            \"dtype\": str(df[col].dtype)\n        })\n\nmissing_df = pd.DataFrame(missing_summary)\n\ndisplay(\n    missing_df\n    .sort_values([\"table\", \"missing_fraction\"], ascending=[True, False])\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:36:12.796998Z","iopub.execute_input":"2026-09-21T02:36:12.797434Z","iopub.status.idle":"2026-09-21T02:36:12.871503Z","shell.execute_reply.started":"2026-09-21T02:36:12.797401Z","shell.execute_reply":"2026-09-21T02:36:12.870122Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Missingness visualization\n# ============================================================\n\nfor name, df in metadata.items():\n\n    missing = df.isna().mean()\n    missing = missing[missing > 0].sort_values(ascending=False)\n\n    if len(missing) == 0:\n        print(f\"{name}: No missing values detected.\")\n        continue\n\n    plt.figure(figsize=(10, 4))\n    missing.plot(kind=\"bar\")\n    plt.ylabel(\"Missing fraction\")\n    plt.xlabel(\"Column\")\n    plt.title(f\"Missingness — {name}\")\n    plt.xticks(rotation=45, ha=\"right\")\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:36:17.744485Z","iopub.execute_input":"2026-09-21T02:36:17.745202Z","iopub.status.idle":"2026-09-21T02:36:18.030591Z","shell.execute_reply.started":"2026-09-21T02:36:17.745159Z","shell.execute_reply":"2026-09-21T02:36:18.027785Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.4) Numeric Summary**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.5 — Numerical summary\n# ============================================================\n\nfor name, df in metadata.items():\n\n    numeric_df = df.select_dtypes(include=np.number)\n\n    print(\"\\n\" + \"=\" * 100)\n    print(f\"{name} — NUMERICAL SUMMARY\")\n    print(\"=\" * 100)\n\n    if numeric_df.shape[1] == 0:\n        print(\"No numerical columns.\")\n        continue\n\n    display(\n        numeric_df.describe(percentiles=[\n            0.01, 0.05, 0.25, 0.50, 0.75, 0.95, 0.99\n        ]).T\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:36:21.775736Z","iopub.execute_input":"2026-09-21T02:36:21.776342Z","iopub.status.idle":"2026-09-21T02:36:22.31987Z","shell.execute_reply.started":"2026-09-21T02:36:21.776303Z","shell.execute_reply":"2026-09-21T02:36:22.318518Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.5) Outlier Screening**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.7 — IQR-based outlier screening\n# ============================================================\n\noutlier_records = []\n\nfor name, df in metadata.items():\n\n    numeric_cols = df.select_dtypes(include=np.number).columns\n\n    for col in numeric_cols:\n\n        values = df[col].dropna()\n\n        if len(values) < 5:\n            continue\n\n        q1 = values.quantile(0.25)\n        q3 = values.quantile(0.75)\n\n        iqr = q3 - q1\n\n        if iqr == 0:\n            outlier_count = 0\n        else:\n            lower = q1 - 1.5 * iqr\n            upper = q3 + 1.5 * iqr\n\n            outlier_count = ((values < lower) | (values > upper)).sum()\n\n        outlier_records.append({\n            \"table\": name,\n            \"column\": col,\n            \"outlier_count\": outlier_count,\n            \"outlier_fraction\": outlier_count / len(values)\n        })\n\noutlier_df = pd.DataFrame(outlier_records)\n\ndisplay(\n    outlier_df\n    .sort_values(\"outlier_fraction\", ascending=False)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:36:50.978979Z","iopub.execute_input":"2026-09-21T02:36:50.979433Z","iopub.status.idle":"2026-09-21T02:36:51.488428Z","shell.execute_reply.started":"2026-09-21T02:36:50.979403Z","shell.execute_reply":"2026-09-21T02:36:51.487242Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.6) Train vs Test Metadata Comaprison**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.8 — Train/Test metadata comparison\n# ============================================================\n\ncommon_columns = sorted(\n    set(train_star_info.columns)\n    & set(test_star_info.columns)\n)\n\nprint(\"Common columns:\")\nprint(common_columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:37:04.817684Z","iopub.execute_input":"2026-09-21T02:37:04.818438Z","iopub.status.idle":"2026-09-21T02:37:04.82516Z","shell.execute_reply.started":"2026-09-21T02:37:04.818399Z","shell.execute_reply":"2026-09-21T02:37:04.82395Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Numerical train/test comparison\n# ============================================================\n\ncommon_numeric = [\n    col for col in common_columns\n    if pd.api.types.is_numeric_dtype(train_star_info[col])\n    and pd.api.types.is_numeric_dtype(test_star_info[col])\n]\n\nfor col in common_numeric:\n\n    train_values = train_star_info[col].dropna()\n    test_values = test_star_info[col].dropna()\n\n    plt.figure(figsize=(8, 4))\n\n    plt.hist(\n        train_values,\n        bins=50,\n        alpha=0.5,\n        label=\"Train\",\n        density=True\n    )\n\n    plt.hist(\n        test_values,\n        bins=50,\n        alpha=0.5,\n        label=\"Test\",\n        density=True\n    )\n\n    plt.xlabel(col)\n    plt.ylabel(\"Density\")\n    plt.title(f\"Train vs Test — {col}\")\n    plt.legend()\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:37:08.713363Z","iopub.execute_input":"2026-09-21T02:37:08.71379Z","iopub.status.idle":"2026-09-21T02:37:11.980017Z","shell.execute_reply.started":"2026-09-21T02:37:08.713755Z","shell.execute_reply":"2026-09-21T02:37:11.97838Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Train/Test numerical summary\n# ============================================================\n\ncomparison_records = []\n\nfor col in common_numeric:\n\n    tr = train_star_info[col]\n    te = test_star_info[col]\n\n    comparison_records.append({\n        \"column\": col,\n\n        \"train_missing\": tr.isna().mean(),\n        \"test_missing\": te.isna().mean(),\n\n        \"train_mean\": tr.mean(),\n        \"test_mean\": te.mean(),\n\n        \"train_std\": tr.std(),\n        \"test_std\": te.std(),\n\n        \"train_median\": tr.median(),\n        \"test_median\": te.median(),\n\n        \"train_min\": tr.min(),\n        \"test_min\": te.min(),\n\n        \"train_max\": tr.max(),\n        \"test_max\": te.max(),\n    })\n\ntrain_test_comparison = pd.DataFrame(comparison_records)\n\ndisplay(train_test_comparison)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:37:19.213163Z","iopub.execute_input":"2026-09-21T02:37:19.213553Z","iopub.status.idle":"2026-09-21T02:37:19.248817Z","shell.execute_reply.started":"2026-09-21T02:37:19.213522Z","shell.execute_reply":"2026-09-21T02:37:19.247249Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.7)Correlation Analysis**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.9 — Correlation analysis\n# ============================================================\n\nfor name, df in metadata.items():\n\n    numeric_df = df.select_dtypes(include=np.number)\n\n    if numeric_df.shape[1] < 2:\n        continue\n\n    corr = numeric_df.corr()\n\n    plt.figure(figsize=(\n        max(8, 0.8 * corr.shape[1]),\n        max(6, 0.8 * corr.shape[0])\n    ))\n\n    plt.imshow(corr, aspect=\"auto\")\n\n    plt.colorbar(label=\"Correlation\")\n\n    plt.xticks(\n        range(len(corr.columns)),\n        corr.columns,\n        rotation=90\n    )\n\n    plt.yticks(\n        range(len(corr.columns)),\n        corr.columns\n    )\n\n    plt.title(f\"Correlation Matrix — {name}\")\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:37:24.202888Z","iopub.execute_input":"2026-09-21T02:37:24.203991Z","iopub.status.idle":"2026-09-21T02:39:00.274505Z","shell.execute_reply.started":"2026-09-21T02:37:24.203946Z","shell.execute_reply":"2026-09-21T02:39:00.273375Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 6.10 — Save metadata audit outputs\n# ============================================================\n\nOUTPUT_DIR = Path(\"/kaggle/working/ariel_eda\")\nOUTPUT_DIR.mkdir(exist_ok=True)\n\nmissing_df.to_parquet(\n    OUTPUT_DIR / \"metadata_missingness.parquet\",\n    index=False\n)\n\noutlier_df.to_parquet(\n    OUTPUT_DIR / \"metadata_outlier_screening.parquet\",\n    index=False\n)\n\ntrain_test_comparison.to_parquet(\n    OUTPUT_DIR / \"train_test_metadata_comparison.parquet\",\n    index=False\n)\n\nprint(\"Saved metadata audit outputs to:\")\nprint(OUTPUT_DIR)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:00.2765Z","iopub.execute_input":"2026-09-21T02:39:00.276936Z","iopub.status.idle":"2026-09-21T02:39:00.29466Z","shell.execute_reply.started":"2026-09-21T02:39:00.2769Z","shell.execute_reply":"2026-09-21T02:39:00.293631Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.8)Target Spectrum EDA**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 6.11 — TARGET SPECTRUM EDA\n# 6.11.1 Target matrix structure\n# ============================================================\n\ntarget_cols = [\n    c for c in train_labels.columns\n    if c.startswith(\"wl_\")\n]\n\nY = train_labels[target_cols].to_numpy()\n\nprint(\"Target matrix shape:\", Y.shape)\nprint(\"Number of planets:\", Y.shape[0])\nprint(\"Number of target channels:\", Y.shape[1])\n\nprint(\"\\nTarget dtype:\", Y.dtype)\nprint(\"Global minimum:\", Y.min())\nprint(\"Global maximum:\", Y.max())\nprint(\"Global mean:\", Y.mean())\nprint(\"Global std:\", Y.std())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:00.296008Z","iopub.execute_input":"2026-09-21T02:39:00.296651Z","iopub.status.idle":"2026-09-21T02:39:00.317087Z","shell.execute_reply.started":"2026-09-21T02:39:00.296611Z","shell.execute_reply":"2026-09-21T02:39:00.315925Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Mean target spectrum\n# ============================================================\n\nmean_spectrum = Y.mean(axis=0)\nstd_spectrum = Y.std(axis=0)\n\nplt.figure(figsize=(12, 5))\n\nplt.plot(\n    np.arange(1, len(target_cols) + 1),\n    mean_spectrum\n)\n\nplt.xlabel(\"Target wavelength channel\")\nplt.ylabel(\"Target value\")\nplt.title(\"Mean Training Spectrum\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:25.401979Z","iopub.execute_input":"2026-09-21T02:39:25.403257Z","iopub.status.idle":"2026-09-21T02:39:25.627558Z","shell.execute_reply.started":"2026-09-21T02:39:25.403205Z","shell.execute_reply":"2026-09-21T02:39:25.626342Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Mean ± 1 STD spectrum\n# ============================================================\n\nx = np.arange(1, len(target_cols) + 1)\n\nplt.figure(figsize=(12, 5))\n\nplt.plot(\n    x,\n    mean_spectrum,\n    label=\"Mean\"\n)\n\nplt.fill_between(\n    x,\n    mean_spectrum - std_spectrum,\n    mean_spectrum + std_spectrum,\n    alpha=0.25,\n    label=\"±1 STD\"\n)\n\nplt.xlabel(\"Target wavelength channel\")\nplt.ylabel(\"Target value\")\nplt.title(\"Training Spectrum — Mean and Cross-Planet Variability\")\nplt.legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:29.757086Z","iopub.execute_input":"2026-09-21T02:39:29.758257Z","iopub.status.idle":"2026-09-21T02:39:29.989961Z","shell.execute_reply.started":"2026-09-21T02:39:29.758215Z","shell.execute_reply":"2026-09-21T02:39:29.988882Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Individual target spectra\n# ============================================================\n\nrng = np.random.default_rng(42)\n\nsample_indices = rng.choice(\n    len(train_labels),\n    size=10,\n    replace=False\n)\n\nplt.figure(figsize=(12, 6))\n\nfor idx in sample_indices:\n\n    plt.plot(\n        x,\n        Y[idx],\n        alpha=0.6\n    )\n\nplt.xlabel(\"Target wavelength channel\")\nplt.ylabel(\"Target value\")\nplt.title(\"Random Sample of Training Spectra\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:35.470154Z","iopub.execute_input":"2026-09-21T02:39:35.470864Z","iopub.status.idle":"2026-09-21T02:39:35.681938Z","shell.execute_reply.started":"2026-09-21T02:39:35.470807Z","shell.execute_reply":"2026-09-21T02:39:35.680648Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Adjacent wavelength correlation\n# ============================================================\n\nadjacent_corr = np.array([\n    np.corrcoef(Y[:, i], Y[:, i + 1])[0, 1]\n    for i in range(Y.shape[1] - 1)\n])\n\nprint(\"Mean adjacent-channel correlation:\",\n      adjacent_corr.mean())\n\nprint(\"Minimum adjacent-channel correlation:\",\n      adjacent_corr.min())\n\nprint(\"Maximum adjacent-channel correlation:\",\n      adjacent_corr.max())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:40.103279Z","iopub.execute_input":"2026-09-21T02:39:40.103769Z","iopub.status.idle":"2026-09-21T02:39:40.143379Z","shell.execute_reply.started":"2026-09-21T02:39:40.103734Z","shell.execute_reply":"2026-09-21T02:39:40.142184Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\n\nplt.plot(\n    np.arange(1, len(adjacent_corr) + 1),\n    adjacent_corr\n)\n\nplt.xlabel(\"Channel pair starting at channel\")\nplt.ylabel(\"Pearson correlation\")\nplt.title(\"Adjacent Target-Channel Correlation\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:44.162649Z","iopub.execute_input":"2026-09-21T02:39:44.16312Z","iopub.status.idle":"2026-09-21T02:39:44.405082Z","shell.execute_reply.started":"2026-09-21T02:39:44.163088Z","shell.execute_reply":"2026-09-21T02:39:44.40394Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Spectral first differences\n# ============================================================\n\nspectral_diff = np.diff(Y, axis=1)\n\nmean_abs_diff = np.mean(\n    np.abs(spectral_diff),\n    axis=0\n)\n\nplt.figure(figsize=(12, 4))\n\nplt.plot(\n    np.arange(1, Y.shape[1]),\n    mean_abs_diff\n)\n\nplt.xlabel(\"Channel\")\nplt.ylabel(\"Mean |Δ spectrum|\")\nplt.title(\"Mean Absolute Adjacent Spectral Difference\")\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:47.637051Z","iopub.execute_input":"2026-09-21T02:39:47.637918Z","iopub.status.idle":"2026-09-21T02:39:48.332663Z","shell.execute_reply.started":"2026-09-21T02:39:47.637882Z","shell.execute_reply":"2026-09-21T02:39:48.331218Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Per-planet spectral statistics\n# ============================================================\n\nplanet_spectral_stats = pd.DataFrame({\n    \"planet_id\": train_labels[\"planet_id\"],\n\n    \"spectrum_mean\": Y.mean(axis=1),\n\n    \"spectrum_std\": Y.std(axis=1),\n\n    \"spectrum_min\": Y.min(axis=1),\n\n    \"spectrum_max\": Y.max(axis=1),\n\n    \"spectrum_range\": Y.max(axis=1) - Y.min(axis=1),\n\n    \"mean_abs_gradient\": np.mean(\n        np.abs(np.diff(Y, axis=1)),\n        axis=1\n    ),\n\n    \"spectral_std_gradient\": np.std(\n        np.diff(Y, axis=1),\n        axis=1\n    )\n})\n\ndisplay(\n    planet_spectral_stats.describe().T\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:51.818355Z","iopub.execute_input":"2026-09-21T02:39:51.819496Z","iopub.status.idle":"2026-09-21T02:39:51.860695Z","shell.execute_reply.started":"2026-09-21T02:39:51.819455Z","shell.execute_reply":"2026-09-21T02:39:51.85946Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Per-planet spectral statistics\n# ============================================================\n\nstats_to_plot = [\n    \"spectrum_mean\",\n    \"spectrum_std\",\n    \"spectrum_range\",\n    \"mean_abs_gradient\",\n    \"spectral_std_gradient\"\n]\n\nfor col in stats_to_plot:\n\n    plt.figure(figsize=(8, 4))\n\n    plt.hist(\n        planet_spectral_stats[col],\n        bins=50\n    )\n\n    plt.xlabel(col)\n    plt.ylabel(\"Number of planets\")\n    plt.title(f\"Distribution — {col}\")\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:39:58.832192Z","iopub.execute_input":"2026-09-21T02:39:58.83268Z","iopub.status.idle":"2026-09-21T02:40:00.084458Z","shell.execute_reply.started":"2026-09-21T02:39:58.832638Z","shell.execute_reply":"2026-09-21T02:40:00.083206Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Spectral anomaly candidates\n# ============================================================\n\nfor col in [\n    \"spectrum_mean\",\n    \"spectrum_std\",\n    \"spectrum_range\",\n    \"mean_abs_gradient\"\n]:\n\n    threshold_low = planet_spectral_stats[col].quantile(0.01)\n    threshold_high = planet_spectral_stats[col].quantile(0.99)\n\n    mask = (\n        (planet_spectral_stats[col] < threshold_low) |\n        (planet_spectral_stats[col] > threshold_high)\n    )\n\n    print(\n        f\"{col}: \"\n        f\"{mask.sum()} planets outside 1st–99th percentile\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:06.652252Z","iopub.execute_input":"2026-09-21T02:40:06.652901Z","iopub.status.idle":"2026-09-21T02:40:06.674322Z","shell.execute_reply.started":"2026-09-21T02:40:06.652802Z","shell.execute_reply":"2026-09-21T02:40:06.67295Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"planet_spectral_stats.to_parquet(\n    OUTPUT_DIR / \"planet_spectral_statistics.parquet\",\n    index=False\n)\n\nprint(\"Saved planet spectral statistics.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:10.740332Z","iopub.execute_input":"2026-09-21T02:40:10.741693Z","iopub.status.idle":"2026-09-21T02:40:10.757383Z","shell.execute_reply.started":"2026-09-21T02:40:10.741643Z","shell.execute_reply":"2026-09-21T02:40:10.756157Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **5.9)Correlate Target Descriptors with Metadata**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 6.12\n# Target ↔ Stellar / Orbital Metadata\n# ============================================================\n\n# Merge spectral descriptors with stellar/orbital metadata\n\nplanet_analysis = (\n    train_star_info\n    .merge(\n        planet_spectral_stats,\n        on=\"planet_id\",\n        how=\"inner\",\n        validate=\"one_to_one\"\n    )\n)\n\nprint(\"Merged shape:\", planet_analysis.shape)\n\ndisplay(planet_analysis.head())\n\nprint(\"\\nMissing values:\")\ndisplay(planet_analysis.isna().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:06:28.887054Z","iopub.execute_input":"2026-09-21T03:06:28.888308Z","iopub.status.idle":"2026-09-21T03:06:28.924116Z","shell.execute_reply.started":"2026-09-21T03:06:28.88826Z","shell.execute_reply":"2026-09-21T03:06:28.923146Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# 6.12.2 Pearson + Spearman correlations\n# ============================================================\n\nmetadata_features = [\n    \"Rs\",\n    \"Ms\",\n    \"Ts\",\n    \"Mp\",\n    \"e\",\n    \"P\",\n    \"sma\",\n    \"i\"\n]\n\nspectral_features = [\n    \"spectrum_mean\",\n    \"spectrum_std\",\n    \"spectrum_min\",\n    \"spectrum_max\",\n    \"spectrum_range\",\n    \"mean_abs_gradient\",\n    \"spectral_std_gradient\"\n]\n\nanalysis_cols = metadata_features + spectral_features\n\npearson_corr = (\n    planet_analysis[analysis_cols]\n    .corr(method=\"pearson\")\n)\n\nspearman_corr = (\n    planet_analysis[analysis_cols]\n    .corr(method=\"spearman\")\n)\n\nprint(\"Pearson correlation:\")\ndisplay(\n    pearson_corr.loc[\n        metadata_features,\n        spectral_features\n    ]\n)\n\nprint(\"\\nSpearman correlation:\")\ndisplay(\n    spearman_corr.loc[\n        metadata_features,\n        spectral_features\n    ]\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:20.07912Z","iopub.execute_input":"2026-09-21T02:40:20.079543Z","iopub.status.idle":"2026-09-21T02:40:20.119981Z","shell.execute_reply.started":"2026-09-21T02:40:20.079511Z","shell.execute_reply":"2026-09-21T02:40:20.118753Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Strongest metadata ↔ spectral relationships\n# ============================================================\n\nrecords = []\n\nfor meta_col in metadata_features:\n\n    for spec_col in spectral_features:\n\n        value = spearman_corr.loc[meta_col, spec_col]\n\n        records.append({\n            \"metadata\": meta_col,\n            \"spectral_feature\": spec_col,\n            \"spearman\": value,\n            \"abs_spearman\": abs(value)\n        })\n\nrelationship_df = (\n    pd.DataFrame(records)\n    .sort_values(\"abs_spearman\", ascending=False)\n)\n\ndisplay(\n    relationship_df.head(15)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:25.009569Z","iopub.execute_input":"2026-09-21T02:40:25.010068Z","iopub.status.idle":"2026-09-21T02:40:25.033156Z","shell.execute_reply.started":"2026-09-21T02:40:25.010032Z","shell.execute_reply":"2026-09-21T02:40:25.031539Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Visualize top relationships\n# ============================================================\n\ntop_relationships = relationship_df.head(6)\n\nfor _, row in top_relationships.iterrows():\n\n    meta_col = row[\"metadata\"]\n    spec_col = row[\"spectral_feature\"]\n\n    plt.figure(figsize=(7, 5))\n\n    plt.scatter(\n        planet_analysis[meta_col],\n        planet_analysis[spec_col],\n        alpha=0.5\n    )\n\n    plt.xlabel(meta_col)\n    plt.ylabel(spec_col)\n\n    plt.title(\n        f\"{spec_col} vs {meta_col}\\n\"\n        f\"Spearman ρ = {row['spearman']:.3f}\"\n    )\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:28.932524Z","iopub.execute_input":"2026-09-21T02:40:28.93314Z","iopub.status.idle":"2026-09-21T02:40:30.155276Z","shell.execute_reply.started":"2026-09-21T02:40:28.933101Z","shell.execute_reply":"2026-09-21T02:40:30.154127Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Spectrum mean vs stellar/orbital parameters\n# ============================================================\n\nplot_features = [\n    \"Rs\",\n    \"Ms\",\n    \"Ts\",\n    \"Mp\",\n    \"P\",\n    \"sma\",\n    \"i\"\n]\n\nfor col in plot_features:\n\n    plt.figure(figsize=(7, 5))\n\n    plt.scatter(\n        planet_analysis[col],\n        planet_analysis[\"spectrum_mean\"],\n        alpha=0.5\n    )\n\n    plt.xlabel(col)\n    plt.ylabel(\"Mean target spectrum\")\n    plt.title(f\"Spectrum Mean vs {col}\")\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-20T19:39:19.532468Z","iopub.execute_input":"2026-09-20T19:39:19.532774Z","iopub.status.idle":"2026-09-20T19:39:20.981604Z","shell.execute_reply.started":"2026-09-20T19:39:19.53275Z","shell.execute_reply":"2026-09-20T19:39:20.980436Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Planet-wise mean normalization\n# ============================================================\n\nY_mean_normalized = (\n    Y /\n    np.mean(Y, axis=1, keepdims=True)\n)\n\nprint(\"Original shape:\", Y.shape)\nprint(\"Normalized shape:\", Y_mean_normalized.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:37.697989Z","iopub.execute_input":"2026-09-21T02:40:37.698519Z","iopub.status.idle":"2026-09-21T02:40:37.708468Z","shell.execute_reply.started":"2026-09-21T02:40:37.698485Z","shell.execute_reply":"2026-09-21T02:40:37.707385Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Adjacent correlation after amplitude normalization\n# ============================================================\n\nnormalized_adjacent_corr = np.array([\n    np.corrcoef(\n        Y_mean_normalized[:, i],\n        Y_mean_normalized[:, i + 1]\n    )[0, 1]\n    for i in range(Y.shape[1] - 1)\n])\n\nprint(\n    \"Mean normalized adjacent correlation:\",\n    normalized_adjacent_corr.mean()\n)\n\nprint(\n    \"Minimum normalized adjacent correlation:\",\n    normalized_adjacent_corr.min()\n)\n\nprint(\n    \"Maximum normalized adjacent correlation:\",\n    normalized_adjacent_corr.max()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:38.910531Z","iopub.execute_input":"2026-09-21T02:40:38.910953Z","iopub.status.idle":"2026-09-21T02:40:38.949005Z","shell.execute_reply.started":"2026-09-21T02:40:38.910922Z","shell.execute_reply":"2026-09-21T02:40:38.947802Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 4))\n\nplt.plot(\n    np.arange(1, len(normalized_adjacent_corr) + 1),\n    normalized_adjacent_corr\n)\n\nplt.xlabel(\"Channel pair starting at channel\")\nplt.ylabel(\"Pearson correlation\")\nplt.title(\n    \"Adjacent Spectral Correlation \"\n    \"After Planet-wise Mean Normalization\"\n)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:42.71296Z","iopub.execute_input":"2026-09-21T02:40:42.713551Z","iopub.status.idle":"2026-09-21T02:40:42.957754Z","shell.execute_reply.started":"2026-09-21T02:40:42.713511Z","shell.execute_reply":"2026-09-21T02:40:42.956633Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# Normalized spectral statistics\n# ============================================================\n\nnormalized_stats = pd.DataFrame({\n    \"planet_id\": train_labels[\"planet_id\"],\n\n    \"normalized_std\":\n        Y_mean_normalized.std(axis=1),\n\n    \"normalized_range\":\n        Y_mean_normalized.max(axis=1)\n        - Y_mean_normalized.min(axis=1),\n\n    \"normalized_mean_abs_gradient\":\n        np.mean(\n            np.abs(np.diff(Y_mean_normalized, axis=1)),\n            axis=1\n        )\n})\n\ndisplay(\n    normalized_stats.describe().T\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:46.623111Z","iopub.execute_input":"2026-09-21T02:40:46.623473Z","iopub.status.idle":"2026-09-21T02:40:46.660484Z","shell.execute_reply.started":"2026-09-21T02:40:46.623434Z","shell.execute_reply":"2026-09-21T02:40:46.65943Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# PCA on normalized spectra\n# ============================================================\n\nfrom sklearn.decomposition import PCA\n\n# Center the normalized spectra across planets\nY_norm_centered = (\n    Y_mean_normalized\n    - Y_mean_normalized.mean(axis=0, keepdims=True)\n)\n\npca = PCA()\n\npca.fit(Y_norm_centered)\n\nexplained_variance = pca.explained_variance_ratio_\ncumulative_variance = np.cumsum(explained_variance)\n\nprint(\n    \"Variance explained by first 5 components:\",\n    explained_variance[:5]\n)\n\nfor threshold in [0.90, 0.95, 0.99]:\n\n    n_components = np.argmax(\n        cumulative_variance >= threshold\n    ) + 1\n\n    print(\n        f\"Components for {threshold:.0%} variance:\",\n        n_components\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:53.941231Z","iopub.execute_input":"2026-09-21T02:40:53.941604Z","iopub.status.idle":"2026-09-21T02:40:55.582269Z","shell.execute_reply.started":"2026-09-21T02:40:53.941576Z","shell.execute_reply":"2026-09-21T02:40:55.581324Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\n\nplt.plot(\n    np.arange(1, len(cumulative_variance) + 1),\n    cumulative_variance\n)\n\nplt.axhline(\n    0.90,\n    linestyle=\"--\",\n    label=\"90%\"\n)\n\nplt.axhline(\n    0.95,\n    linestyle=\"--\",\n    label=\"95%\"\n)\n\nplt.axhline(\n    0.99,\n    linestyle=\"--\",\n    label=\"99%\"\n)\n\nplt.xlabel(\"Number of PCA components\")\nplt.ylabel(\"Cumulative explained variance\")\nplt.title(\"Spectral Low-Rank Structure\")\n\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:40:59.669897Z","iopub.execute_input":"2026-09-21T02:40:59.671295Z","iopub.status.idle":"2026-09-21T02:40:59.91949Z","shell.execute_reply.started":"2026-09-21T02:40:59.671251Z","shell.execute_reply":"2026-09-21T02:40:59.91831Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ============================================================\n# First-channel discontinuity diagnostic\n# ============================================================\n\nfirst_diff = np.abs(Y[:, 1] - Y[:, 0])\n\nremaining_diff = np.abs(np.diff(Y[:, 1:], axis=1))\n\nprint(\"First-channel mean absolute difference:\")\nprint(first_diff.mean())\n\nprint(\"\\nRemaining-channel mean absolute difference:\")\nprint(remaining_diff.mean())\n\nprint(\"\\nRatio:\")\nprint(first_diff.mean() / remaining_diff.mean())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:41:03.622612Z","iopub.execute_input":"2026-09-21T02:41:03.623588Z","iopub.status.idle":"2026-09-21T02:41:03.638388Z","shell.execute_reply.started":"2026-09-21T02:41:03.623549Z","shell.execute_reply":"2026-09-21T02:41:03.636935Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **6)Raw Detector Signal EDA**","metadata":{}},{"cell_type":"markdown","source":"## **6.1)Sample detector File**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.1 — RAW SIGNAL FILE INSPECTION\n# ============================================================\n\nfrom pathlib import Path\nimport pandas as pd\nimport numpy as np\n\nARIEL_ROOT = Path(\"/kaggle/input/competitions/ariel-data-challenge-2025\")\nTRAIN_ROOT = ARIEL_ROOT / \"train\"\nTEST_ROOT = ARIEL_ROOT / \"test\"\n\n# ------------------------------------------------------------\n# Select one training planet\n# ------------------------------------------------------------\n\nsample_planet = train_planets[0]\n\nsample_dir = TRAIN_ROOT / sample_planet\n\nprint(\"=\" * 80)\nprint(\"SAMPLE PLANET\")\nprint(\"=\" * 80)\nprint(\"Planet ID:\", sample_planet)\nprint(\"Directory:\", sample_dir)\n\nsignal_files = sorted(sample_dir.glob(\"*signal*.parquet\"))\n\nprint(\"\\nSignal files:\")\nfor f in signal_files:\n    print(\" \", f.name)\n\n# ------------------------------------------------------------\n# Inspect parquet metadata without loading full data\n# ------------------------------------------------------------\n\nfor signal_file in signal_files:\n\n    print(\"\\n\" + \"=\" * 80)\n    print(signal_file.name)\n    print(\"=\" * 80)\n\n    df = pd.read_parquet(signal_file)\n\n    print(\"Shape:\", df.shape)\n    print(\"Dtypes:\")\n    print(df.dtypes.value_counts())\n\n    print(\"\\nFirst 5 rows:\")\n    display(df.head())\n\n    print(\"\\nMemory usage:\")\n    print(f\"{df.memory_usage(deep=True).sum() / 1024**2:.2f} MB\")\n\n    print(\"\\nNaN count:\", df.isna().sum().sum())\n    print(\"Inf count:\", np.isinf(df.select_dtypes(include=np.number)).sum().sum())\n\n    print(\"\\nBasic statistics:\")\n    display(\n        df.describe(\n            percentiles=[0.001, 0.01, 0.5, 0.99, 0.999]\n        )\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:53:24.593156Z","iopub.execute_input":"2026-09-21T02:53:24.594823Z","iopub.status.idle":"2026-09-21T02:54:03.440529Z","shell.execute_reply.started":"2026-09-21T02:53:24.594774Z","shell.execute_reply":"2026-09-21T02:54:03.439284Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.2) Detector Geometry**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.2 — DETECTOR GEOMETRY\n# ============================================================\n\nprint(\"=\" * 80)\nprint(\"AXIS INFORMATION\")\nprint(\"=\" * 80)\n\nprint(axis_info.shape)\ndisplay(axis_info.head(10))\n\nprint(\"\\nColumns:\")\nfor col in axis_info.columns:\n    print(f\"  {col}: {axis_info[col].shape}\")\n\n# ------------------------------------------------------------\n# Unique integration times\n# ------------------------------------------------------------\n\nprint(\"\\nAIRS integration times:\")\nprint(\n    np.sort(\n        axis_info[\"AIRS-CH0-integration_time\"]\n        .dropna()\n        .unique()\n    )\n)\n\n# ------------------------------------------------------------\n# Axis ranges\n# ------------------------------------------------------------\n\nfor col in axis_info.columns:\n\n    values = axis_info[col].dropna()\n\n    print(\n        f\"{col:35s} \"\n        f\"n={len(values):6d} \"\n        f\"min={values.min():.8f} \"\n        f\"max={values.max():.8f}\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:54:51.274511Z","iopub.execute_input":"2026-09-21T02:54:51.274964Z","iopub.status.idle":"2026-09-21T02:54:51.303476Z","shell.execute_reply.started":"2026-09-21T02:54:51.274932Z","shell.execute_reply":"2026-09-21T02:54:51.302297Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.3) Temporal Statistics**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.3 — TEMPORAL STATISTICS\n# ============================================================\n\ndef temporal_statistics(df, name):\n\n    arr = df.to_numpy(dtype=np.float64)\n\n    # Mean over detector/spectral dimension\n    frame_mean = np.nanmean(arr, axis=1)\n\n    # Standard deviation over detector/spectral dimension\n    frame_std = np.nanstd(arr, axis=1)\n\n    # Median\n    frame_median = np.nanmedian(arr, axis=1)\n\n    # Robust temporal statistics\n    q01 = np.nanpercentile(arr, 1, axis=1)\n    q99 = np.nanpercentile(arr, 99, axis=1)\n\n    stats = pd.DataFrame({\n        \"frame_mean\": frame_mean,\n        \"frame_std\": frame_std,\n        \"frame_median\": frame_median,\n        \"frame_q01\": q01,\n        \"frame_q99\": q99,\n    })\n\n    print(\"=\" * 80)\n    print(name)\n    print(\"=\" * 80)\n\n    display(stats.describe())\n\n    return stats\n\n\n# Example:\nairs_file = sorted(sample_dir.glob(\"AIRS-CH0_signal_*.parquet\"))[0]\nfgs_file = sorted(sample_dir.glob(\"FGS1_signal_*.parquet\"))[0]\n\nairs_df = pd.read_parquet(airs_file)\nfgs_df = pd.read_parquet(fgs_file)\n\nairs_temporal = temporal_statistics(\n    airs_df,\n    \"AIRS temporal statistics\"\n)\n\nfgs_temporal = temporal_statistics(\n    fgs_df,\n    \"FGS1 temporal statistics\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:55:28.52427Z","iopub.execute_input":"2026-09-21T02:55:28.524684Z","iopub.status.idle":"2026-09-21T02:56:23.764081Z","shell.execute_reply.started":"2026-09-21T02:55:28.524654Z","shell.execute_reply":"2026-09-21T02:56:23.762758Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.4) Temporal Detector Behaviour**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.4 — TEMPORAL DETECTOR BEHAVIOR\n# ============================================================\n\nfig, axes = plt.subplots(3, 1, figsize=(14, 12))\n\naxes[0].plot(airs_temporal[\"frame_mean\"])\naxes[0].set_title(\"AIRS — Frame Mean\")\naxes[0].set_xlabel(\"Frame\")\naxes[0].set_ylabel(\"Mean Signal\")\naxes[0].grid(True)\n\naxes[1].plot(airs_temporal[\"frame_std\"])\naxes[1].set_title(\"AIRS — Frame Standard Deviation\")\naxes[1].set_xlabel(\"Frame\")\naxes[1].set_ylabel(\"STD\")\naxes[1].grid(True)\n\naxes[2].plot(airs_temporal[\"frame_median\"])\naxes[2].set_title(\"AIRS — Frame Median\")\naxes[2].set_xlabel(\"Frame\")\naxes[2].set_ylabel(\"Median Signal\")\naxes[2].grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:56:23.765958Z","iopub.execute_input":"2026-09-21T02:56:23.766537Z","iopub.status.idle":"2026-09-21T02:56:28.296016Z","shell.execute_reply.started":"2026-09-21T02:56:23.766487Z","shell.execute_reply":"2026-09-21T02:56:28.294653Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.5) Spectral Profile**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.5 — MEAN DETECTOR / SPECTRAL PROFILE\n# ============================================================\n\nairs_array = airs_df.to_numpy(dtype=np.float64)\n\nmean_profile = np.nanmean(airs_array, axis=0)\nstd_profile = np.nanstd(airs_array, axis=0)\n\nfig, ax = plt.subplots(figsize=(15, 5))\n\nax.plot(mean_profile)\n\nax.set_title(\"AIRS — Mean Signal Across Detector Channels\")\nax.set_xlabel(\"Detector / Spectral Channel\")\nax.set_ylabel(\"Mean Signal\")\nax.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:57:14.04071Z","iopub.execute_input":"2026-09-21T02:57:14.042178Z","iopub.status.idle":"2026-09-21T02:57:17.689935Z","shell.execute_reply.started":"2026-09-21T02:57:14.042112Z","shell.execute_reply":"2026-09-21T02:57:17.68873Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 5))\n\nax.plot(std_profile)\n\nax.set_title(\"AIRS — Temporal Variability per Detector Channel\")\nax.set_xlabel(\"Detector / Spectral Channel\")\nax.set_ylabel(\"Temporal STD\")\nax.grid(True)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:57:30.36903Z","iopub.execute_input":"2026-09-21T02:57:30.369482Z","iopub.status.idle":"2026-09-21T02:57:30.60799Z","shell.execute_reply.started":"2026-09-21T02:57:30.369447Z","shell.execute_reply":"2026-09-21T02:57:30.606828Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.6) Extreme Channel Identification**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.6 — EXTREME CHANNEL DIAGNOSTICS\n# ============================================================\n\nchannel_stats = pd.DataFrame({\n    \"channel\": np.arange(len(mean_profile)),\n    \"mean\": mean_profile,\n    \"std\": std_profile,\n})\n\nchannel_stats[\"cv\"] = (\n    channel_stats[\"std\"] /\n    np.abs(channel_stats[\"mean\"])\n)\n\nprint(\"Highest mean-signal channels:\")\ndisplay(\n    channel_stats\n    .sort_values(\"mean\", ascending=False)\n    .head(10)\n)\n\nprint(\"\\nHighest variability channels:\")\ndisplay(\n    channel_stats\n    .sort_values(\"std\", ascending=False)\n    .head(10)\n)\n\nprint(\"\\nHighest coefficient-of-variation channels:\")\ndisplay(\n    channel_stats\n    .sort_values(\"cv\", ascending=False)\n    .head(10)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:59:18.702949Z","iopub.execute_input":"2026-09-21T02:59:18.70412Z","iopub.status.idle":"2026-09-21T02:59:18.744227Z","shell.execute_reply.started":"2026-09-21T02:59:18.704073Z","shell.execute_reply":"2026-09-21T02:59:18.742996Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.7)Outlier Frame Detection**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.7 — TEMPORAL OUTLIER DETECTION\n# ============================================================\n\nframe_stats = airs_temporal.copy()\n\n# Robust z-score using median absolute deviation\ndef robust_zscore(x):\n    x = np.asarray(x)\n    med = np.nanmedian(x)\n    mad = np.nanmedian(np.abs(x - med))\n\n    if mad == 0:\n        return np.zeros_like(x)\n\n    return 0.6745 * (x - med) / mad\n\n\nframe_stats[\"robust_z_mean\"] = robust_zscore(\n    frame_stats[\"frame_mean\"]\n)\n\nframe_stats[\"robust_z_std\"] = robust_zscore(\n    frame_stats[\"frame_std\"]\n)\n\noutlier_frames = frame_stats[\n    (np.abs(frame_stats[\"robust_z_mean\"]) > 5) |\n    (np.abs(frame_stats[\"robust_z_std\"]) > 5)\n]\n\nprint(\"Number of potential outlier frames:\", len(outlier_frames))\n\ndisplay(outlier_frames.head(20))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T02:59:45.897734Z","iopub.execute_input":"2026-09-21T02:59:45.89852Z","iopub.status.idle":"2026-09-21T02:59:45.918391Z","shell.execute_reply.started":"2026-09-21T02:59:45.89848Z","shell.execute_reply":"2026-09-21T02:59:45.917125Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **6.8) Compact Raw-Data Diagnostics**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.8 — SAVE COMPACT RAW-DATA DIAGNOSTICS\n# ============================================================\n\ndiagnostic_dir = Path(\"/kaggle/working/ariel_diagnostics\")\ndiagnostic_dir.mkdir(exist_ok=True)\n\nchannel_stats.to_parquet(\n    diagnostic_dir / f\"{sample_planet}_AIRS_channel_stats.parquet\",\n    index=False\n)\n\nframe_stats.to_parquet(\n    diagnostic_dir / f\"{sample_planet}_AIRS_frame_stats.parquet\",\n    index=False\n)\n\nprint(\"Saved compact diagnostics to:\")\nprint(diagnostic_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:00:30.782714Z","iopub.execute_input":"2026-09-21T03:00:30.783789Z","iopub.status.idle":"2026-09-21T03:00:30.845435Z","shell.execute_reply.started":"2026-09-21T03:00:30.783736Z","shell.execute_reply":"2026-09-21T03:00:30.84426Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **7) Raw Detector EDA**","metadata":{}},{"cell_type":"markdown","source":"## **7.1) Representative Planet Selection**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 6 — METADATA + TARGET SPECTRUM INTEGRATION\n# Robust version\n# ============================================================\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\n# ------------------------------------------------------------\n# 1. Identify target wavelength columns\n# ------------------------------------------------------------\n\ntarget_cols = [\n    c for c in train_labels.columns\n    if c.startswith(\"wl_\")\n]\n\nprint(f\"Number of target wavelength channels: {len(target_cols)}\")\nprint(f\"First 5: {target_cols[:5]}\")\nprint(f\"Last 5:  {target_cols[-5:]}\")\n\n\n# ------------------------------------------------------------\n# 2. Extract target spectra\n# ------------------------------------------------------------\n\ntarget_matrix = train_labels[target_cols].to_numpy(dtype=np.float64)\n\nprint(\"Target matrix shape:\", target_matrix.shape)\n\n\n# ------------------------------------------------------------\n# 3. Compute spectral statistics\n# ------------------------------------------------------------\n\nspectral_stats = pd.DataFrame({\n    \"planet_id\": train_labels[\"planet_id\"].astype(np.int64),\n\n    \"spectrum_mean\":\n        np.mean(target_matrix, axis=1),\n\n    \"spectrum_std\":\n        np.std(target_matrix, axis=1),\n\n    \"spectrum_min\":\n        np.min(target_matrix, axis=1),\n\n    \"spectrum_max\":\n        np.max(target_matrix, axis=1),\n\n    \"spectrum_range\":\n        np.ptp(target_matrix, axis=1),\n\n    \"mean_abs_gradient\":\n        np.mean(np.abs(np.diff(target_matrix, axis=1)), axis=1),\n\n    \"spectral_std_gradient\":\n        np.std(np.diff(target_matrix, axis=1), axis=1)\n})\n\nprint(\"\\nSpectral statistics:\")\ndisplay(spectral_stats.head())\n\n\n# ------------------------------------------------------------\n# 4. Make planet_id type consistent\n# ------------------------------------------------------------\n\n# train_manifest was created from directory names,\n# therefore planet_id may currently be string/object.\n\ntrain_manifest = train_manifest.copy()\n\ntrain_manifest[\"planet_id\"] = pd.to_numeric(\n    train_manifest[\"planet_id\"],\n    errors=\"coerce\"\n).astype(\"Int64\")\n\nspectral_stats[\"planet_id\"] = pd.to_numeric(\n    spectral_stats[\"planet_id\"],\n    errors=\"coerce\"\n).astype(\"Int64\")\n\n\n# ------------------------------------------------------------\n# 5. Verify ID compatibility before merge\n# ------------------------------------------------------------\n\nmanifest_ids = set(train_manifest[\"planet_id\"].dropna())\ntarget_ids = set(spectral_stats[\"planet_id\"].dropna())\n\nprint(\"\\nID audit\")\nprint(\"-\" * 60)\nprint(\"Manifest planets :\", len(manifest_ids))\nprint(\"Target planets   :\", len(target_ids))\nprint(\"Common IDs       :\", len(manifest_ids & target_ids))\nprint(\"Only manifest    :\", len(manifest_ids - target_ids))\nprint(\"Only target      :\", len(target_ids - manifest_ids))\n\n\n# ------------------------------------------------------------\n# 6. Merge manifest + spectral statistics\n# ------------------------------------------------------------\n\nselection_df = train_manifest.merge(\n    spectral_stats,\n    on=\"planet_id\",\n    how=\"inner\",\n    validate=\"one_to_one\"\n)\n\nprint(\"\\nMerged shape:\", selection_df.shape)\n\ndisplay(selection_df.head())\n\n\n# ------------------------------------------------------------\n# 7. Add stellar metadata\n# ------------------------------------------------------------\n\ntrain_star_clean = train_star_info.copy()\n\ntrain_star_clean[\"planet_id\"] = pd.to_numeric(\n    train_star_clean[\"planet_id\"],\n    errors=\"coerce\"\n).astype(\"Int64\")\n\nselection_df = selection_df.merge(\n    train_star_clean,\n    on=\"planet_id\",\n    how=\"left\",\n    validate=\"one_to_one\"\n)\n\nprint(\"\\nAfter adding stellar metadata:\")\nprint(\"Shape:\", selection_df.shape)\n\ndisplay(selection_df.head())\n\n\n# ------------------------------------------------------------\n# 8. Missing-value audit\n# ------------------------------------------------------------\n\nmissing_summary = (\n    selection_df.isna()\n    .sum()\n    .sort_values(ascending=False)\n)\n\nprint(\"\\nMissing values:\")\ndisplay(\n    missing_summary[missing_summary > 0]\n)\n\n\n# ------------------------------------------------------------\n# 9. Final feature list\n# ------------------------------------------------------------\n\nspectral_features = [\n    \"spectrum_mean\",\n    \"spectrum_std\",\n    \"spectrum_min\",\n    \"spectrum_max\",\n    \"spectrum_range\",\n    \"mean_abs_gradient\",\n    \"spectral_std_gradient\"\n]\n\nstellar_features = [\n    \"Rs\",\n    \"Ms\",\n    \"Ts\",\n    \"Mp\",\n    \"e\",\n    \"P\",\n    \"sma\",\n    \"i\"\n]\n\nprint(\"\\nFinal spectral features:\")\nprint(spectral_features)\n\nprint(\"\\nFinal stellar features:\")\nprint(stellar_features)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:12:02.69336Z","iopub.execute_input":"2026-09-21T03:12:02.693758Z","iopub.status.idle":"2026-09-21T03:12:02.808393Z","shell.execute_reply.started":"2026-09-21T03:12:02.69373Z","shell.execute_reply":"2026-09-21T03:12:02.807195Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **7.2) Raw Signal File Structure**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.2 — RAW SIGNAL FILE STRUCTURE\n# ============================================================\n\nfrom pathlib import Path\nimport pyarrow.parquet as pq\n\n# ------------------------------------------------------------\n# 1. Select representative planets\n# ------------------------------------------------------------\n# Use the previously generated representative_planets.\n# If the variable does not exist, use the following fallback.\n\nif \"representative_planets\" not in globals():\n\n    representative_planets = (\n        selection_df\n        .sort_values(\"spectrum_mean\")\n        .iloc[\n            np.linspace(\n                0,\n                len(selection_df) - 1,\n                5\n            ).astype(int)\n        ][\"planet_id\"]\n        .astype(int)\n        .tolist()\n    )\n\nprint(\"Representative planets:\")\nprint(representative_planets)\n\n\n# ------------------------------------------------------------\n# 2. Inspect directory structure\n# ------------------------------------------------------------\n\nfor planet_id in representative_planets:\n\n    planet_dir = TRAIN_ROOT / str(planet_id)\n\n    print(\"\\n\" + \"=\" * 100)\n    print(f\"PLANET: {planet_id}\")\n    print(f\"DIRECTORY: {planet_dir}\")\n    print(\"=\" * 100)\n\n    if not planet_dir.exists():\n        print(\"WARNING: directory does not exist\")\n        continue\n\n    files = sorted(\n        [\n            f for f in planet_dir.rglob(\"*\")\n            if f.is_file()\n        ]\n    )\n\n    print(f\"Total files: {len(files)}\")\n\n    for f in files:\n        print(\" \", f.relative_to(planet_dir))\n\n\n# ------------------------------------------------------------\n# 3. Inspect Parquet metadata\n# ------------------------------------------------------------\n\ndef inspect_parquet_metadata(path):\n\n    path = Path(path)\n\n    print(\"\\n\" + \"-\" * 100)\n    print(f\"FILE: {path.name}\")\n    print(f\"SIZE: {path.stat().st_size / (1024**2):.2f} MB\")\n    print(\"-\" * 100)\n\n    pf = pq.ParquetFile(path)\n\n    print(\"Rows       :\", pf.metadata.num_rows)\n    print(\"Columns    :\", pf.metadata.num_columns)\n    print(\"Row groups :\", pf.num_row_groups)\n\n    print(\"\\nColumns:\")\n    for i, column in enumerate(\n        pf.schema_arrow.names\n    ):\n        print(f\"  {i:3d}: {column}\")\n\n    print(\"\\nArrow schema:\")\n    print(pf.schema_arrow)\n\n\n# ------------------------------------------------------------\n# 4. Get signal files for first representative planet\n# ------------------------------------------------------------\n\nsample_planet = int(representative_planets[0])\n\nsample_dir = TRAIN_ROOT / str(sample_planet)\n\nsignal_files = sorted(\n    [\n        f\n        for f in sample_dir.rglob(\"*signal*.parquet\")\n        if f.is_file()\n    ]\n)\n\nprint(\"\\n\" + \"=\" * 100)\nprint(f\"SIGNAL FILES FOR PLANET {sample_planet}\")\nprint(\"=\" * 100)\n\nfor f in signal_files:\n    print(f)\n\n\n# ------------------------------------------------------------\n# 5. Separate instruments\n# ------------------------------------------------------------\n\nairs_files = [\n    f for f in signal_files\n    if \"AIRS\" in f.name\n]\n\nfgs_files = [\n    f for f in signal_files\n    if \"FGS1\" in f.name\n]\n\nprint(\"\\nAIRS signal files:\", len(airs_files))\nprint(\"FGS1 signal files:\", len(fgs_files))\n\n\n# ------------------------------------------------------------\n# 6. Inspect one file from each instrument\n# ------------------------------------------------------------\n\nif len(airs_files) > 0:\n\n    print(\"\\n\\nAIRS FILE\")\n    inspect_parquet_metadata(\n        airs_files[0]\n    )\n\n\nif len(fgs_files) > 0:\n\n    print(\"\\n\\nFGS1 FILE\")\n    inspect_parquet_metadata(\n        fgs_files[0]\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:14:15.15637Z","iopub.execute_input":"2026-09-21T03:14:15.158025Z","iopub.status.idle":"2026-09-21T03:14:15.926782Z","shell.execute_reply.started":"2026-09-21T03:14:15.157909Z","shell.execute_reply":"2026-09-21T03:14:15.925269Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **7.3) Raw Detector Geometry**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# SECTION 7.3 — RAW DETECTOR GEOMETRY\n# ============================================================\n\nimport numpy as np\nimport pandas as pd\nimport pyarrow.parquet as pq\nimport matplotlib.pyplot as plt\n\n# ------------------------------------------------------------\n# 1. Select one representative planet\n# ------------------------------------------------------------\n\nsample_planet = int(representative_planets[0])\n\nsample_dir = TRAIN_ROOT / str(sample_planet)\n\nairs_path = sample_dir / \"AIRS-CH0_signal_0.parquet\"\nfgs_path  = sample_dir / \"FGS1_signal_0.parquet\"\n\nprint(\"Planet:\", sample_planet)\nprint(\"AIRS:\", airs_path)\nprint(\"FGS1 :\", fgs_path)\n\n\n# ------------------------------------------------------------\n# 2. Memory-safe row reader\n# ------------------------------------------------------------\n\ndef read_parquet_rows(path, start=0, stop=100):\n    \"\"\"\n    Read only a small row range from a Parquet file.\n\n    Returns a NumPy array.\n    \"\"\"\n    \n    pf = pq.ParquetFile(path)\n    \n    stop = min(stop, pf.metadata.num_rows)\n    \n    table = pf.read(\n        columns=None,\n        use_threads=True\n    )\n    \n    # Convert only the requested rows after reading the table.\n    # This is acceptable for this small diagnostic only if\n    # the file is manageable, but we will avoid this approach\n    # for the full dataset.\n    df = table.slice(start, stop - start).to_pandas()\n    \n    return df.to_numpy()\n\n\n# ------------------------------------------------------------\n# IMPORTANT:\n# Because the AIRS file is ~104 MB, first inspect only\n# metadata and then read a limited set of columns.\n# ------------------------------------------------------------\n\ndef read_selected_columns(path, row_start=0, row_stop=20,\n                          col_indices=None):\n    \n    pf = pq.ParquetFile(path)\n    \n    if col_indices is None:\n        col_indices = list(range(min(20, pf.metadata.num_columns)))\n    \n    column_names = [\n        pf.schema.names[i]\n        for i in col_indices\n    ]\n    \n    table = pf.read(\n        columns=column_names,\n        use_threads=True\n    )\n    \n    table = table.slice(\n        row_start,\n        min(row_stop, table.num_rows) - row_start\n    )\n    \n    return table.to_pandas()\n\n\n# ------------------------------------------------------------\n# 3. Read a small AIRS detector patch\n# ------------------------------------------------------------\n\nairs_patch = read_selected_columns(\n    airs_path,\n    row_start=0,\n    row_stop=20,\n    col_indices=np.linspace(\n        0,\n        pq.ParquetFile(airs_path).metadata.num_columns - 1,\n        20,\n        dtype=int\n    )\n)\n\nprint(\"\\nAIRS selected detector patch:\")\ndisplay(airs_patch)\n\nprint(\"\\nAIRS patch shape:\", airs_patch.shape)\nprint(\"AIRS patch dtype:\")\nprint(airs_patch.dtypes.value_counts())\n\n\n# ------------------------------------------------------------\n# 4. Read a small FGS1 detector patch\n# ------------------------------------------------------------\n\nfgs_patch = read_selected_columns(\n    fgs_path,\n    row_start=0,\n    row_stop=20,\n    col_indices=np.linspace(\n        0,\n        pq.ParquetFile(fgs_path).metadata.num_columns - 1,\n        20,\n        dtype=int\n    )\n)\n\nprint(\"\\nFGS1 selected detector patch:\")\ndisplay(fgs_patch)\n\nprint(\"\\nFGS1 patch shape:\", fgs_patch.shape)\nprint(\"FGS1 patch dtype:\")\nprint(fgs_patch.dtypes.value_counts())\n\n\n# ------------------------------------------------------------\n# 5. Detector-column statistics\n# ------------------------------------------------------------\n\ndef detector_column_stats(path, n_rows=100):\n    \n    pf = pq.ParquetFile(path)\n    \n    n_cols = pf.metadata.num_columns\n    \n    # Read first n_rows across all columns.\n    # This gives us temporal variation across the detector\n    # without loading the entire file.\n    table = pf.read(\n        columns=pf.schema.names,\n        use_threads=True\n    ).slice(0, min(n_rows, pf.metadata.num_rows))\n    \n    arr = table.to_pandas().to_numpy(dtype=np.float32)\n    \n    stats = pd.DataFrame({\n        \"column\": np.arange(n_cols),\n        \"mean\": np.mean(arr, axis=0),\n        \"std\": np.std(arr, axis=0),\n        \"min\": np.min(arr, axis=0),\n        \"max\": np.max(arr, axis=0),\n        \"range\": np.ptp(arr, axis=0)\n    })\n    \n    return stats\n\n\nprint(\"\\nComputing AIRS detector statistics...\")\nairs_detector_stats = detector_column_stats(\n    airs_path,\n    n_rows=100\n)\n\nprint(airs_detector_stats.head())\nprint(airs_detector_stats.tail())\n\n\nprint(\"\\nComputing FGS1 detector statistics...\")\nfgs_detector_stats = detector_column_stats(\n    fgs_path,\n    n_rows=100\n)\n\nprint(fgs_detector_stats.head())\nprint(fgs_detector_stats.tail())\n\n\n# ------------------------------------------------------------\n# 6. Detector mean profiles\n# ------------------------------------------------------------\n\nfig, ax = plt.subplots(figsize=(14, 4))\n\nax.plot(\n    airs_detector_stats[\"column\"],\n    airs_detector_stats[\"mean\"],\n    linewidth=1\n)\n\nax.set_title(\n    f\"AIRS raw detector mean profile — Planet {sample_planet}\"\n)\n\nax.set_xlabel(\"Detector column\")\nax.set_ylabel(\"Raw uint16 mean\")\n\nplt.tight_layout()\nplt.show()\n\n\nfig, ax = plt.subplots(figsize=(14, 4))\n\nax.plot(\n    fgs_detector_stats[\"column\"],\n    fgs_detector_stats[\"mean\"],\n    linewidth=1\n)\n\nax.set_title(\n    f\"FGS1 raw detector mean profile — Planet {sample_planet}\"\n)\n\nax.set_xlabel(\"Detector column\")\nax.set_ylabel(\"Raw uint16 mean\")\n\nplt.tight_layout()\nplt.show()\n\n\n# ------------------------------------------------------------\n# 7. Temporal variation of selected detector pixels\n# ------------------------------------------------------------\n\ndef read_selected_pixel_timeseries(path, pixel_indices,\n                                   n_rows=500):\n    \n    pf = pq.ParquetFile(path)\n    \n    column_names = [\n        pf.schema.names[i]\n        for i in pixel_indices\n    ]\n    \n    table = pf.read(\n        columns=column_names,\n        use_threads=True\n    )\n    \n    table = table.slice(\n        0,\n        min(n_rows, table.num_rows)\n    )\n    \n    return table.to_pandas()\n\n\nairs_pixels = [\n    0,\n    100,\n    500,\n    1000,\n    2000,\n    4000,\n    6000,\n    8000,\n    10000,\n    11391\n]\n\nairs_ts = read_selected_pixel_timeseries(\n    airs_path,\n    airs_pixels,\n    n_rows=500\n)\n\nprint(\"\\nAIRS selected-pixel time series:\")\ndisplay(airs_ts.head())\n\n\nfig, ax = plt.subplots(figsize=(14, 5))\n\nfor col in airs_ts.columns:\n    ax.plot(\n        airs_ts.index,\n        airs_ts[col],\n        linewidth=1,\n        label=col\n    )\n\nax.set_title(\n    f\"AIRS raw temporal variation — Planet {sample_planet}\"\n)\n\nax.set_xlabel(\"Raw integration index\")\nax.set_ylabel(\"Raw detector value\")\n\nax.legend(\n    ncol=2,\n    fontsize=8\n)\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:18:37.899644Z","iopub.execute_input":"2026-09-21T03:18:37.900103Z","iopub.status.idle":"2026-09-21T03:18:45.529385Z","shell.execute_reply.started":"2026-09-21T03:18:37.900064Z","shell.execute_reply":"2026-09-21T03:18:45.527895Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **7.4) Decode Raw Integration**","metadata":{}},{"cell_type":"code","source":"# ============================================================\n# 7.4 — Decode Raw Integration / Readout Structure\n# ============================================================\n#\n# Objective:\n#   1. Understand how axis_info maps to AIRS / FGS1.\n#   2. Identify valid AIRS integrations.\n#   3. Characterize integration-time alternation.\n#   4. Recover the AIRS physical time axis.\n#   5. Validate raw signal ↔ metadata alignment.\n#   6. Characterize AIRS wavelength metadata.\n#   7. Test the 32 × 356 AIRS geometry hypothesis.\n#   8. Test the 32 × 32 FGS1 geometry hypothesis.\n#\n# We deliberately DO NOT reshape the detector yet.\n# The flattening convention is investigated in Section 7.5.\n# ============================================================\n\n\n# ------------------------------------------------------------\n# 0. Imports and representative file paths\n# ------------------------------------------------------------\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport pyarrow.parquet as pq\n\n# Use the same representative planet selected in Section 7.3\nsample_planet = int(representative_planets[0])\n\nsample_dir = TRAIN_ROOT / str(sample_planet)\n\nairs_path = sample_dir / \"AIRS-CH0_signal_0.parquet\"\nfgs_path  = sample_dir / \"FGS1_signal_0.parquet\"\n\nprint(\"=\" * 70)\nprint(\"SECTION 7.4 — RAW INTEGRATION / READOUT STRUCTURE\")\nprint(\"=\" * 70)\n\nprint(\"\\nRepresentative planet:\", sample_planet)\nprint(\"AIRS:\", airs_path)\nprint(\"FGS1 :\", fgs_path)\n\n\n# ------------------------------------------------------------\n# 1. Load axis metadata\n# ------------------------------------------------------------\n\naxis_path = ARIEL_ROOT / \"axis_info.parquet\"\n\naxis_info_local = pd.read_parquet(axis_path)\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"1. AXIS METADATA OVERVIEW\")\nprint(\"=\" * 70)\n\nprint(\"\\naxis_info shape:\", axis_info_local.shape)\n\nprint(\"\\nColumns:\")\nfor col in axis_info_local.columns:\n    print(\" -\", col)\n\nprint(\"\\nFirst 10 rows:\")\ndisplay(axis_info_local.head(10))\n\nprint(\"\\nDtypes:\")\ndisplay(axis_info_local.dtypes)\n\nprint(\"\\nUnique-value / missing-value counts:\")\n\naxis_summary = pd.DataFrame({\n    \"dtype\": axis_info_local.dtypes.astype(str),\n    \"n_unique\": axis_info_local.nunique(dropna=True),\n    \"n_missing\": axis_info_local.isna().sum(),\n    \"missing_fraction\": axis_info_local.isna().mean()\n})\n\ndisplay(axis_summary)\n\n\n# ------------------------------------------------------------\n# 2. Separate AIRS and FGS1 metadata\n# ------------------------------------------------------------\n\nairs_axis_cols = [\n    c for c in axis_info_local.columns\n    if c.startswith(\"AIRS-CH0\")\n]\n\nfgs_axis_cols = [\n    c for c in axis_info_local.columns\n    if c.startswith(\"FGS1\")\n]\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"2. INSTRUMENT-SPECIFIC METADATA\")\nprint(\"=\" * 70)\n\nprint(\"\\nAIRS metadata columns:\")\nfor col in airs_axis_cols:\n    print(\" -\", col)\n\nprint(\"\\nFGS1 metadata columns:\")\nfor col in fgs_axis_cols:\n    print(\" -\", col)\n\n\n# ------------------------------------------------------------\n# 3. Extract VALID AIRS metadata\n# ------------------------------------------------------------\n#\n# axis_info is a combined table.\n#\n# AIRS fields occupy only the AIRS portion of the table;\n# the remainder is NaN.\n#\n# Therefore we MUST NOT use:\n#\n#     axis_info_local.iloc[:135000]\n#\n# as AIRS metadata.\n#\n# Instead, identify rows where the AIRS time coordinate exists.\n# ------------------------------------------------------------\n\nairs_axis = axis_info_local[\n    axis_info_local[\"AIRS-CH0-axis0-h\"].notna()\n].copy()\n\nfgs_axis = axis_info_local[\n    axis_info_local[\"FGS1-axis0-h\"].notna()\n].copy()\n\nprint(\"\\nValid metadata lengths:\")\nprint(\"AIRS:\", len(airs_axis))\nprint(\"FGS1:\", len(fgs_axis))\n\nprint(\"\\nAIRS missing values after filtering:\")\ndisplay(airs_axis.isna().sum())\n\nprint(\"\\nFGS1 missing values:\")\ndisplay(fgs_axis.isna().sum())\n\n\n# ------------------------------------------------------------\n# 4. Inspect AIRS metadata\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"3. AIRS METADATA\")\nprint(\"=\" * 70)\n\ndisplay(airs_axis.head(20))\n\nprint(\"\\nAIRS unique values:\")\n\nfor col in airs_axis.columns:\n\n    print(f\"\\n{col}\")\n\n    print(\n        \"  unique:\",\n        airs_axis[col].nunique(dropna=True)\n    )\n\n    print(\n        \"  missing:\",\n        airs_axis[col].isna().sum()\n    )\n\n    print(\n        \"  first values:\",\n        airs_axis[col]\n        .dropna()\n        .drop_duplicates()\n        .head(10)\n        .to_list()\n    )\n\n\n# ------------------------------------------------------------\n# 5. Inspect FGS1 metadata\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"4. FGS1 METADATA\")\nprint(\"=\" * 70)\n\ndisplay(fgs_axis.head(20))\n\nprint(\"\\nFGS1 axis0-h statistics:\")\n\nprint(\n    fgs_axis[\"FGS1-axis0-h\"].describe()\n)\n\n\n# ------------------------------------------------------------\n# 6. Read raw Parquet metadata\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"5. RAW FILE DIMENSIONS\")\nprint(\"=\" * 70)\n\nairs_pf = pq.ParquetFile(airs_path)\nfgs_pf = pq.ParquetFile(fgs_path)\n\nairs_rows = airs_pf.metadata.num_rows\nairs_columns = airs_pf.metadata.num_columns\n\nfgs_rows = fgs_pf.metadata.num_rows\nfgs_columns = fgs_pf.metadata.num_columns\n\nprint(\"\\nAIRS raw:\")\nprint(\"  rows   :\", airs_rows)\nprint(\"  columns:\", airs_columns)\n\nprint(\"\\nFGS1 raw:\")\nprint(\"  rows   :\", fgs_rows)\nprint(\"  columns:\", fgs_columns)\n\n\n# ------------------------------------------------------------\n# 7. Verify metadata ↔ raw dimensions\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"6. METADATA / RAW DIMENSION CONSISTENCY\")\nprint(\"=\" * 70)\n\nairs_metadata_length = len(airs_axis)\nfgs_metadata_length = len(fgs_axis)\n\nprint(\n    \"\\nAIRS metadata rows:\",\n    airs_metadata_length\n)\n\nprint(\n    \"AIRS raw rows:\",\n    airs_rows\n)\n\nprint(\n    \"AIRS row match:\",\n    airs_metadata_length == airs_rows\n)\n\nprint(\n    \"\\nFGS1 metadata rows:\",\n    fgs_metadata_length\n)\n\nprint(\n    \"FGS1 raw rows:\",\n    fgs_rows\n)\n\nprint(\n    \"FGS1 row match:\",\n    fgs_metadata_length == fgs_rows\n)\n\nassert airs_metadata_length == airs_rows, (\n    \"AIRS metadata length does not match raw AIRS rows.\"\n)\n\nassert fgs_metadata_length == fgs_rows, (\n    \"FGS1 metadata length does not match raw FGS1 rows.\"\n)\n\n\n# ------------------------------------------------------------\n# 8. AIRS integration-time distribution\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"7. AIRS INTEGRATION-TIME STRUCTURE\")\nprint(\"=\" * 70)\n\nairs_time_col = \"AIRS-CH0-integration_time\"\n\nintegration_values = (\n    airs_axis[airs_time_col]\n    .to_numpy()\n)\n\nprint(\"\\nIntegration-time value counts:\")\n\nintegration_counts = (\n    pd.Series(integration_values)\n    .value_counts(dropna=False)\n    .sort_index()\n)\n\ndisplay(integration_counts)\n\nprint(\"\\nUnique integration times:\")\nprint(\n    np.unique(\n        integration_values[\n            ~pd.isna(integration_values)\n        ]\n    )\n)\n\n\n# ------------------------------------------------------------\n# 9. Examine first 100 integration times\n# ------------------------------------------------------------\n\nprint(\"\\nFirst 100 AIRS integration times:\")\n\nprint(\n    integration_values[:100]\n)\n\n\n# ------------------------------------------------------------\n# 10. Test whether integration times alternate\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"8. INTEGRATION-TIME ALTERNATION TEST\")\nprint(\"=\" * 70)\n\nexpected_pattern = np.resize(\n    np.array([0.1, 4.5]),\n    len(integration_values)\n)\n\nvalid_integration_mask = ~pd.isna(integration_values)\n\nalternation_matches = (\n    integration_values[valid_integration_mask]\n    ==\n    expected_pattern[valid_integration_mask]\n)\n\nprint(\n    \"Rows matching expected [0.1, 4.5] pattern:\",\n    alternation_matches.sum()\n)\n\nprint(\n    \"Total valid AIRS rows:\",\n    valid_integration_mask.sum()\n)\n\nprint(\n    \"Fraction matching:\",\n    alternation_matches.mean()\n)\n\nprint(\n    \"\\nExact alternating pattern:\",\n    bool(alternation_matches.all())\n)\n\n\n# ------------------------------------------------------------\n# 11. Transition analysis\n# ------------------------------------------------------------\n\ntransitions = np.flatnonzero(\n    integration_values[1:]\n    !=\n    integration_values[:-1]\n) + 1\n\nprint(\n    \"\\nNumber of integration-time transitions:\",\n    len(transitions)\n)\n\nprint(\n    \"Expected transitions for strict alternation:\",\n    len(integration_values) - 1\n)\n\nprint(\"\\nFirst 30 transition indices:\")\nprint(transitions[:30])\n\nprint(\"\\nFirst 20 transition details:\")\n\nfor idx in transitions[:20]:\n\n    print(\n        f\"row {idx-1:6d} -> {idx:6d} : \"\n        f\"{integration_values[idx-1]} -> \"\n        f\"{integration_values[idx]}\"\n    )\n\n\n# ------------------------------------------------------------\n# 12. Run-length encoding of integration times\n# ------------------------------------------------------------\n\nruns = []\n\nstart = 0\n\nfor i in range(1, len(integration_values)):\n\n    if integration_values[i] != integration_values[i - 1]:\n\n        runs.append({\n            \"start_row\": start,\n            \"end_row\": i - 1,\n            \"length\": i - start,\n            \"integration_time\": integration_values[i - 1]\n        })\n\n        start = i\n\n\nruns.append({\n    \"start_row\": start,\n    \"end_row\": len(integration_values) - 1,\n    \"length\": len(integration_values) - start,\n    \"integration_time\": integration_values[-1]\n})\n\nairs_runs = pd.DataFrame(runs)\n\nprint(\"\\nIntegration-time run structure:\")\ndisplay(airs_runs.head(30))\n\nprint(\"\\nRun-length statistics:\")\n\ndisplay(\n    airs_runs\n    .groupby(\"integration_time\")[\"length\"]\n    .describe()\n)\n\n\n# ------------------------------------------------------------\n# 13. AIRS time coordinate\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"9. AIRS PHYSICAL TIME AXIS\")\nprint(\"=\" * 70)\n\nairs_h_col = \"AIRS-CH0-axis0-h\"\n\nairs_h = (\n    airs_axis[airs_h_col]\n    .to_numpy()\n)\n\nprint(\"\\naxis0-h length:\", len(airs_h))\n\nprint(\"dtype:\", airs_h.dtype)\n\nprint(\"first 20:\")\nprint(airs_h[:20])\n\nprint(\"\\nTime-axis statistics:\")\n\nprint(\n    \"min:\",\n    np.nanmin(airs_h)\n)\n\nprint(\n    \"max:\",\n    np.nanmax(airs_h)\n)\n\n\n# ------------------------------------------------------------\n# 14. Convert axis0-h from hours to seconds\n# ------------------------------------------------------------\n\nairs_time_seconds = airs_h * 3600.0\n\nprint(\"\\nTime in seconds:\")\n\nprint(\n    \"start:\",\n    airs_time_seconds[0]\n)\n\nprint(\n    \"end:\",\n    airs_time_seconds[-1]\n)\n\nprint(\n    \"duration:\",\n    airs_time_seconds[-1]\n    - airs_time_seconds[0],\n    \"seconds\"\n)\n\n\n# ------------------------------------------------------------\n# 15. Time-axis increments\n# ------------------------------------------------------------\n\nairs_h_diff = np.diff(airs_h)\n\nvalid_h_diff = airs_h_diff[\n    ~np.isnan(airs_h_diff)\n]\n\nprint(\"\\naxis0-h differences:\")\n\nprint(\n    \"minimum:\",\n    np.min(valid_h_diff)\n)\n\nprint(\n    \"maximum:\",\n    np.max(valid_h_diff)\n)\n\nprint(\n    \"median:\",\n    np.median(valid_h_diff)\n)\n\nprint(\"\\nMost common axis0-h increments:\")\n\ndisplay(\n    pd.Series(valid_h_diff)\n    .round(12)\n    .value_counts()\n    .head(15)\n)\n\n\n# ------------------------------------------------------------\n# 16. Convert common increments to seconds\n# ------------------------------------------------------------\n\nincrement_counts = (\n    pd.Series(valid_h_diff)\n    .round(12)\n    .value_counts()\n    .head(10)\n)\n\nincrement_seconds = (\n    increment_counts.index.to_numpy()\n    * 3600\n)\n\nincrement_table = pd.DataFrame({\n    \"increment_hours\": increment_counts.index,\n    \"increment_seconds\": increment_seconds,\n    \"count\": increment_counts.values\n})\n\nprint(\"\\nCommon time increments in seconds:\")\n\ndisplay(increment_table)\n\n\n# ------------------------------------------------------------\n# 17. Plot physical time coordinate\n# ------------------------------------------------------------\n\nfig, ax = plt.subplots(figsize=(12, 4))\n\nax.plot(\n    np.arange(len(airs_time_seconds)),\n    airs_time_seconds,\n    linewidth=0.8\n)\n\nax.set_title(\n    f\"AIRS physical time coordinate — Planet {sample_planet}\"\n)\n\nax.set_xlabel(\"Raw AIRS row / integration index\")\n\nax.set_ylabel(\"Time since reference [s]\")\n\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\n\n# ------------------------------------------------------------\n# 18. Plot integration type versus raw row\n# ------------------------------------------------------------\n\nfig, ax = plt.subplots(figsize=(12, 3))\n\nax.plot(\n    np.arange(len(integration_values)),\n    integration_values,\n    linewidth=0.6\n)\n\nax.set_title(\n    \"AIRS integration time versus raw row index\"\n)\n\nax.set_xlabel(\"Raw AIRS row index\")\n\nax.set_ylabel(\"Integration time [s]\")\n\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\n\n# ------------------------------------------------------------\n# 19. Read diagnostic AIRS detector columns\n# ------------------------------------------------------------\n#\n# IMPORTANT:\n# Only selected columns are read.\n# The complete 104 MB AIRS file is NOT loaded.\n# ------------------------------------------------------------\n\ndiagnostic_columns = [\n    5000,\n    5500,\n    6000,\n    6500,\n    7000,\n]\n\ndiagnostic_column_names = [\n    f\"column_{c}\"\n    for c in diagnostic_columns\n]\n\ndiagnostic_table = airs_pf.read(\n    columns=diagnostic_column_names\n)\n\ndiagnostic_df = diagnostic_table.to_pandas()\n\nprint(\"\\nDiagnostic raw signal shape:\")\nprint(diagnostic_df.shape)\n\nprint(\"\\nExpected rows:\")\nprint(airs_rows)\n\nassert len(diagnostic_df) == airs_rows\n\n\n# ------------------------------------------------------------\n# 20. Display diagnostic raw signal\n# ------------------------------------------------------------\n\nprint(\"\\nFirst 20 rows:\")\n\ndisplay(\n    diagnostic_df.head(20)\n)\n\n\n# ------------------------------------------------------------\n# 21. Align diagnostic signal with AIRS metadata\n# ------------------------------------------------------------\n\ntarget_col = 6000\n\ncomparison_full = pd.DataFrame({\n    \"raw_row\": np.arange(len(diagnostic_df)),\n\n    \"raw_value\": (\n        diagnostic_df[f\"column_{target_col}\"]\n        .to_numpy()\n    ),\n\n    \"integration_time_s\": integration_values,\n\n    \"axis0_h\": airs_h,\n\n    \"time_seconds\": airs_time_seconds\n})\n\nprint(\"\\nAligned raw signal + metadata:\")\n\ndisplay(\n    comparison_full.head(30)\n)\n\n\n# ------------------------------------------------------------\n# 22. Verify alignment dimensions\n# ------------------------------------------------------------\n\nassert len(comparison_full) == airs_rows\n\nassert (\n    len(comparison_full[\"raw_value\"])\n    ==\n    len(comparison_full[\"integration_time_s\"])\n)\n\nassert (\n    len(comparison_full[\"raw_value\"])\n    ==\n    len(comparison_full[\"time_seconds\"])\n)\n\nprint(\n    \"\\nRaw signal and AIRS metadata are dimensionally aligned.\"\n)\n\n\n# ------------------------------------------------------------\n# 23. Raw signal statistics by integration time\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"10. RAW SIGNAL VS INTEGRATION TIME\")\nprint(\"=\" * 70)\n\ngrouped_stats = (\n    comparison_full\n    .groupby(\"integration_time_s\")[\"raw_value\"]\n    .agg(\n        count=\"count\",\n        mean=\"mean\",\n        std=\"std\",\n        min=\"min\",\n        median=\"median\",\n        max=\"max\"\n    )\n)\n\ndisplay(grouped_stats)\n\n\n# ------------------------------------------------------------\n# 24. Signal ratio between long and short integrations\n# ------------------------------------------------------------\n\nmeans = grouped_stats[\"mean\"]\n\nif 0.1 in means.index and 4.5 in means.index:\n\n    ratio_4p5_to_0p1 = (\n        means.loc[4.5]\n        /\n        means.loc[0.1]\n    )\n\n    print(\n        \"\\nMean raw signal ratio \"\n        \"(4.5 s / 0.1 s):\",\n        ratio_4p5_to_0p1\n    )\n\n    print(\n        \"\\nIntegration-time ratio:\",\n        4.5 / 0.1\n    )\n\n\n# ------------------------------------------------------------\n# 25. Plot raw signal for first 500 rows\n# ------------------------------------------------------------\n\nfig, ax = plt.subplots(figsize=(13, 4))\n\nax.plot(\n    comparison_full[\"time_seconds\"].iloc[:500],\n    comparison_full[\"raw_value\"].iloc[:500],\n    linewidth=0.8\n)\n\nax.set_title(\n    f\"AIRS raw detector signal vs physical time\\n\"\n    f\"Detector column {target_col}\"\n)\n\nax.set_xlabel(\"Time since observation reference [s]\")\n\nax.set_ylabel(\"Raw detector value\")\n\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\n\n# ------------------------------------------------------------\n# 26. Separate short and long integrations\n# ------------------------------------------------------------\n\nmask_short = (\n    comparison_full[\"integration_time_s\"] == 0.1\n)\n\nmask_long = (\n    comparison_full[\"integration_time_s\"] == 4.5\n)\n\nplot_n = min(500, len(comparison_full))\n\nplot_subset = comparison_full.iloc[:plot_n]\n\nmask_short_plot = (\n    plot_subset[\"integration_time_s\"] == 0.1\n)\n\nmask_long_plot = (\n    plot_subset[\"integration_time_s\"] == 4.5\n)\n\nfig, ax = plt.subplots(figsize=(13, 4))\n\nax.scatter(\n    plot_subset.loc[\n        mask_short_plot,\n        \"time_seconds\"\n    ],\n    plot_subset.loc[\n        mask_short_plot,\n        \"raw_value\"\n    ],\n    s=8,\n    label=\"0.1 s integration\"\n)\n\nax.scatter(\n    plot_subset.loc[\n        mask_long_plot,\n        \"time_seconds\"\n    ],\n    plot_subset.loc[\n        mask_long_plot,\n        \"raw_value\"\n    ],\n    s=8,\n    label=\"4.5 s integration\"\n)\n\nax.set_title(\n    f\"AIRS raw signal separated by integration time\\n\"\n    f\"Detector column {target_col}\"\n)\n\nax.set_xlabel(\"Time [s]\")\n\nax.set_ylabel(\"Raw detector value\")\n\nax.legend()\n\nax.grid(True, alpha=0.3)\n\nplt.tight_layout()\nplt.show()\n\n\n# ------------------------------------------------------------\n# 27. Quantify alternating signal difference\n# ------------------------------------------------------------\n\nshort_values = (\n    comparison_full\n    .loc[mask_short, \"raw_value\"]\n    .to_numpy()\n    .astype(float)\n)\n\nlong_values = (\n    comparison_full\n    .loc[mask_long, \"raw_value\"]\n    .to_numpy()\n    .astype(float)\n)\n\npaired_n = min(\n    len(short_values),\n    len(long_values)\n)\n\nshort_values = short_values[:paired_n]\nlong_values = long_values[:paired_n]\n\npaired_difference = (\n    long_values\n    -\n    short_values\n)\n\npaired_ratio = (\n    long_values\n    /\n    np.maximum(short_values, 1e-12)\n)\n\nprint(\"\\nPaired short/long integration diagnostics:\")\nprint(\n    \"Pairs:\",\n    paired_n\n)\n\nprint(\n    \"Mean difference:\",\n    paired_difference.mean()\n)\n\nprint(\n    \"Median difference:\",\n    np.median(paired_difference)\n)\n\nprint(\n    \"Mean ratio:\",\n    paired_ratio.mean()\n)\n\nprint(\n    \"Median ratio:\",\n    np.median(paired_ratio)\n)\n\n\n# ------------------------------------------------------------\n# 28. AIRS wavelength metadata\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"11. AIRS WAVELENGTH AXIS\")\nprint(\"=\" * 70)\n\nairs_wl_col = \"AIRS-CH0-axis2-um\"\n\nairs_wavelength = (\n    axis_info_local[airs_wl_col]\n    .dropna()\n    .to_numpy()\n)\n\nprint(\n    \"Number of wavelength samples:\",\n    len(airs_wavelength)\n)\n\nprint(\n    \"Number of unique wavelengths:\",\n    np.unique(airs_wavelength).size\n)\n\nprint(\n    \"Minimum wavelength:\",\n    np.min(airs_wavelength),\n    \"µm\"\n)\n\nprint(\n    \"Maximum wavelength:\",\n    np.max(airs_wavelength),\n    \"µm\"\n)\n\nprint(\"\\nFirst 20 wavelengths:\")\nprint(airs_wavelength[:20])\n\nprint(\"\\nLast 20 wavelengths:\")\nprint(airs_wavelength[-20:])\n\n\n# ------------------------------------------------------------\n# 29. Wavelength ordering\n# ------------------------------------------------------------\n\nwl_diff = np.diff(airs_wavelength)\n\nprint(\"\\nWavelength ordering:\")\n\nprint(\n    \"Strictly increasing:\",\n    np.all(wl_diff > 0)\n)\n\nprint(\n    \"Strictly decreasing:\",\n    np.all(wl_diff < 0)\n)\n\nprint(\n    \"Monotonic increasing:\",\n    np.all(wl_diff >= 0)\n)\n\nprint(\n    \"Monotonic decreasing:\",\n    np.all(wl_diff <= 0)\n)\n\n\n# ------------------------------------------------------------\n# 30. Wavelength spacing statistics\n# ------------------------------------------------------------\n\nprint(\"\\nWavelength spacing:\")\n\nprint(\n    \"minimum absolute spacing:\",\n    np.min(np.abs(wl_diff))\n)\n\nprint(\n    \"maximum absolute spacing:\",\n    np.max(np.abs(wl_diff))\n)\n\nprint(\n    \"median absolute spacing:\",\n    np.median(np.abs(wl_diff))\n)\n\nprint(\n    \"mean absolute spacing:\",\n    np.mean(np.abs(wl_diff))\n)\n\n\n# ------------------------------------------------------------\n# 31. Check wavelength repetition\n# ------------------------------------------------------------\n\nwavelength_counts = (\n    pd.Series(airs_wavelength)\n    .value_counts()\n)\n\nprint(\"\\nWavelength repetition statistics:\")\n\nprint(\n    \"minimum count:\",\n    wavelength_counts.min()\n)\n\nprint(\n    \"maximum count:\",\n    wavelength_counts.max()\n)\n\nprint(\n    \"median count:\",\n    wavelength_counts.median()\n)\n\nprint(\"\\nRepetition-count distribution:\")\n\ndisplay(\n    wavelength_counts\n    .value_counts()\n    .sort_index()\n    .head(20)\n)\n\n\n# ------------------------------------------------------------\n# 32. Critical AIRS detector geometry calculation\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"12. AIRS DETECTOR GEOMETRY HYPOTHESIS\")\nprint(\"=\" * 70)\n\nAIRS_RAW_COLUMNS = airs_columns\nAIRS_WAVELENGTH_COUNT = len(airs_wavelength)\n\nprint(\n    \"\\nAIRS raw detector columns:\",\n    AIRS_RAW_COLUMNS\n)\n\nprint(\n    \"AIRS wavelength samples:\",\n    AIRS_WAVELENGTH_COUNT\n)\n\nprint(\n    \"Raw columns / wavelength samples:\",\n    AIRS_RAW_COLUMNS\n    /\n    AIRS_WAVELENGTH_COUNT\n)\n\nprint(\n    \"\\n32 × wavelength count:\",\n    32 * AIRS_WAVELENGTH_COUNT\n)\n\nprint(\n    \"Exact equality:\",\n    AIRS_RAW_COLUMNS\n    ==\n    32 * AIRS_WAVELENGTH_COUNT\n)\n\n\n# ------------------------------------------------------------\n# 33. Explicit geometry check\n# ------------------------------------------------------------\n\nairs_spatial_dimension = (\n    AIRS_RAW_COLUMNS\n    //\n    AIRS_WAVELENGTH_COUNT\n)\n\nairs_geometry_remainder = (\n    AIRS_RAW_COLUMNS\n    %\n    AIRS_WAVELENGTH_COUNT\n)\n\nprint(\n    \"\\nInferred spatial dimension:\",\n    airs_spatial_dimension\n)\n\nprint(\n    \"Remainder:\",\n    airs_geometry_remainder\n)\n\nif airs_geometry_remainder == 0:\n\n    print(\n        \"\\nAIRS detector columns are exactly divisible \"\n        \"by the wavelength-axis length.\"\n    )\n\n    print(\n        f\"Candidate detector geometry: \"\n        f\"{airs_spatial_dimension} × \"\n        f\"{AIRS_WAVELENGTH_COUNT}\"\n    )\n\nelse:\n\n    print(\n        \"\\nAIRS columns are NOT exactly divisible \"\n        \"by the wavelength-axis length.\"\n    )\n\n\n# ------------------------------------------------------------\n# 34. FGS1 detector geometry\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"13. FGS1 DETECTOR GEOMETRY\")\nprint(\"=\" * 70)\n\nFGS_RAW_COLUMNS = fgs_columns\n\nprint(\n    \"FGS1 raw detector columns:\",\n    FGS_RAW_COLUMNS\n)\n\nprint(\n    \"Square-root:\",\n    np.sqrt(FGS_RAW_COLUMNS)\n)\n\nprint(\n    \"32 × 32:\",\n    32 * 32\n)\n\nprint(\n    \"Exactly 32 × 32:\",\n    FGS_RAW_COLUMNS == 32 * 32\n)\n\n\n# ------------------------------------------------------------\n# 35. Do NOT reshape yet\n# ------------------------------------------------------------\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"14. INTERPRETATION STATUS\")\nprint(\"=\" * 70)\n\nprint(\n    \"\"\"\nCurrent evidence:\n\nAIRS\n----\nRaw rows:\n    11,250\n\nRaw detector columns:\n    11,392\n\nAIRS wavelength metadata:\n    356 samples\n\n11,392 = 32 × 356\n\nTherefore:\n    A 32 × 356 detector geometry is strongly\n    supported by dimensional consistency.\n\nHowever:\n    The flattening order of the 32 × 356 detector\n    has NOT yet been established.\n\nFGS1\n----\nRaw rows:\n    135,000\n\nRaw detector columns:\n    1,024\n\n1,024 = 32 × 32\n\nTherefore:\n    A 32 × 32 detector geometry is consistent\n    with the raw dimensionality.\n\nIntegration structure\n---------------------\nAIRS integration_time alternates:\n\n    0.1 s\n    4.5 s\n    0.1 s\n    4.5 s\n    ...\n\nThe raw detector signal shows corresponding\nalternation between low and high signal states.\n\nTherefore:\n    This alternating behavior should be treated\n    as acquisition structure, not removed as an\n    arbitrary temporal artifact.\n\nNext step\n---------\nSection 7.5 will determine the actual flattening\nconvention and reconstruct the 2-D detector geometry.\n\"\"\"\n)\n\n\n# ------------------------------------------------------------\n# 36. Compact machine-readable summary\n# ------------------------------------------------------------\n\nsection_74_summary = {\n    \"planet_id\": sample_planet,\n\n    \"airs_raw_rows\": int(airs_rows),\n    \"airs_raw_columns\": int(airs_columns),\n\n    \"airs_metadata_rows\": int(len(airs_axis)),\n\n    \"airs_wavelength_count\": int(\n        len(airs_wavelength)\n    ),\n\n    \"airs_candidate_spatial_dimension\": int(\n        airs_spatial_dimension\n    ),\n\n    \"airs_geometry_remainder\": int(\n        airs_geometry_remainder\n    ),\n\n    \"airs_integration_times\": sorted(\n        pd.Series(integration_values)\n        .dropna()\n        .unique()\n        .tolist()\n    ),\n\n    \"airs_integration_alternates\": bool(\n        alternation_matches.all()\n    ),\n\n    \"airs_time_start_seconds\": float(\n        airs_time_seconds[0]\n    ),\n\n    \"airs_time_end_seconds\": float(\n        airs_time_seconds[-1]\n    ),\n\n    \"airs_wavelength_min_um\": float(\n        np.min(airs_wavelength)\n    ),\n\n    \"airs_wavelength_max_um\": float(\n        np.max(airs_wavelength)\n    ),\n\n    \"fgs_raw_rows\": int(fgs_rows),\n    \"fgs_raw_columns\": int(fgs_columns),\n\n    \"fgs_candidate_geometry\": \"32 x 32\"\n}\n\nsection_74_summary_df = pd.DataFrame(\n    [section_74_summary]\n).T\n\nsection_74_summary_df.columns = [\n    \"value\"\n]\n\nprint(\"\\n\" + \"=\" * 70)\nprint(\"15. SECTION 7.4 MACHINE-READABLE SUMMARY\")\nprint(\"=\" * 70)\n\ndisplay(section_74_summary_df)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-21T03:30:58.435978Z","iopub.execute_input":"2026-09-21T03:30:58.437565Z","iopub.status.idle":"2026-09-21T03:31:01.058789Z","shell.execute_reply.started":"2026-09-21T03:30:58.437482Z","shell.execute_reply":"2026-09-21T03:31:01.057715Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}