{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.x"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118765,"databundleVersionId":15231210,"sourceType":"competition"}],"dockerImageVersionId":31259,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"e98bc4e7-a25a-4620-af27-d585f96a9b7b","cell_type":"markdown","source":"# Stanford RNA 3D Folding — Starter Notebook (Template)\n\nThis notebook is a **directly runnable Kaggle template** for the Stanford RNA 3D Folding competition.\n\n✅ Out of the box\n- Discovers input files in `/kaggle/input`\n- Loads `train.csv`, `test.csv`, `sample_submission.csv` **if present**\n- Runs a **dummy baseline** that outputs correctly-shaped coordinates\n- Builds `submission.csv`\n\n🧩 Customize\n- Dataset folder name (or let the notebook auto-detect)\n- Column names (`sequence`, `id`, etc.)\n- `predict_coords()` with your real model\n- Submission formatting (some competitions require special column layouts)\n\n> **Leakage warning:** keep external structural data **<= May 29, 2025** (per host guidance). Log external sources.\n","metadata":{}},{"id":"093723f4-77cb-48e5-9866-3de690e71ee7","cell_type":"code","source":"import os, gc, json, math, random\nimport numpy as np\nimport pandas as pd\n\nSEED = 42\nrandom.seed(SEED)\nnp.random.seed(SEED)\n\npd.set_option(\"display.max_columns\", 200)\npd.set_option(\"display.max_rows\", 200)\n\nprint(\"ready\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:17:29.772413Z","iopub.execute_input":"2026-02-02T15:17:29.773322Z","iopub.status.idle":"2026-02-02T15:17:30.124175Z","shell.execute_reply.started":"2026-02-02T15:17:29.773279Z","shell.execute_reply":"2026-02-02T15:17:30.123242Z"}},"outputs":[],"execution_count":null},{"id":"7f5a3955-81dd-464f-8f32-2e7ab7b5ac7c","cell_type":"markdown","source":"## 1) Find dataset folder + list files\n\nIf you already know the dataset directory, set `DATA_DIR` manually.\nOtherwise, this cell lists `/kaggle/input` and tries to auto-detect a folder that contains `sample_submission.csv`.\n","metadata":{}},{"id":"f0175cc6-5c0f-40e6-a3bc-2d358e9f90cf","cell_type":"code","source":"import os\n\nINPUT_ROOT = \"/kaggle/input\"\n\ndef list_input_tree(root=INPUT_ROOT, max_files_per_dir=25):\n    if not os.path.exists(root):\n        print(\"Not running on Kaggle? /kaggle/input not found.\")\n        return\n    for d in sorted(os.listdir(root)):\n        p = os.path.join(root, d)\n        if os.path.isdir(p):\n            print(f\"\\n== {d} ==\")\n            for r, _, files in os.walk(p):\n                rel = r.replace(p, \"\").lstrip(\"/\")\n                if files:\n                    print(f\"  /{rel}\" if rel else \"  /\")\n                    for f in sorted(files)[:max_files_per_dir]:\n                        print(\"   -\", f)\n\ndef autodetect_data_dir(root=INPUT_ROOT):\n    if not os.path.exists(root):\n        return None\n    for d in sorted(os.listdir(root)):\n        p = os.path.join(root, d)\n        if not os.path.isdir(p):\n            continue\n        for r, _, files in os.walk(p):\n            if \"sample_submission.csv\" in files:\n                return p\n    return None\n\nlist_input_tree()\n\nDATA_DIR = autodetect_data_dir()\nprint(\"\\nAuto-detected DATA_DIR:\", DATA_DIR)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:17:55.320027Z","iopub.execute_input":"2026-02-02T15:17:55.320929Z","iopub.status.idle":"2026-02-02T15:18:11.933394Z","shell.execute_reply.started":"2026-02-02T15:17:55.32089Z","shell.execute_reply":"2026-02-02T15:18:11.932232Z"}},"outputs":[],"execution_count":null},{"id":"b354fff9-0f76-45bc-b4ca-52d0b0325a41","cell_type":"markdown","source":"## 2) Load train/test/sample_submission\n\nAssumes common filenames:\n- `train.csv`\n- `test.csv`\n- `sample_submission.csv`\n\nIf your competition uses different names or nested paths, edit below.\n","metadata":{}},{"id":"2af2b66d-f08c-4107-9719-14f119f35b8a","cell_type":"code","source":"import os\nimport pandas as pd\n\ndef safe_read_csv(path):\n    if path and os.path.exists(path):\n        return pd.read_csv(path)\n    return None\n\ntrain_path = os.path.join(DATA_DIR, \"train.csv\") if DATA_DIR else None\ntest_path  = os.path.join(DATA_DIR, \"test.csv\") if DATA_DIR else None\nsub_path   = os.path.join(DATA_DIR, \"sample_submission.csv\") if DATA_DIR else None\n\ntrain = safe_read_csv(train_path)\ntest  = safe_read_csv(test_path)\nsub   = safe_read_csv(sub_path)\n\nprint(\"train:\", None if train is None else train.shape, train_path)\nprint(\"test :\", None if test is None else test.shape, test_path)\nprint(\"sub  :\", None if sub is None else sub.shape, sub_path)\n\nif train is not None:\n    display(train.head())\nif test is not None:\n    display(test.head())\nif sub is not None:\n    display(sub.head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:18:22.451947Z","iopub.execute_input":"2026-02-02T15:18:22.452349Z","iopub.status.idle":"2026-02-02T15:18:22.535633Z","shell.execute_reply.started":"2026-02-02T15:18:22.45232Z","shell.execute_reply":"2026-02-02T15:18:22.534592Z"}},"outputs":[],"execution_count":null},{"id":"8b5d21a1-881f-47a5-a7d0-56bc81ef1be3","cell_type":"markdown","source":"## 3) Set column names\n\nEdit these once.\n\nCommon:\n- ID column: `id` / `target_id` / `sequence_id`\n- sequence column: `sequence`\n","metadata":{}},{"id":"92d23a5d-b566-4500-a7dc-ba7b5f9b51a1","cell_type":"code","source":"# Inspect columns\nif train is not None:\n    print(\"train columns:\", train.columns.tolist())\nif test is not None:\n    print(\"test columns :\", test.columns.tolist())\nif sub is not None:\n    print(\"sub columns  :\", sub.columns.tolist())\n\n# ---- TODO: adjust if needed ----\nID_COL  = \"id\"\nSEQ_COL = \"sequence\"\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:19:05.756013Z","iopub.execute_input":"2026-02-02T15:19:05.75647Z","iopub.status.idle":"2026-02-02T15:19:05.763346Z","shell.execute_reply.started":"2026-02-02T15:19:05.756439Z","shell.execute_reply":"2026-02-02T15:19:05.762093Z"}},"outputs":[],"execution_count":null},{"id":"c23dfefe-8ebb-475c-bd46-ee6d82d69025","cell_type":"markdown","source":"## 4) Quick EDA (sequence lengths)\n","metadata":{}},{"id":"5d69aca5-27d1-4c08-8bc2-2258f0d18479","cell_type":"code","source":"def len_stats(df, col):\n    lens = df[col].astype(str).str.len()\n    return pd.Series({\n        \"n\": len(lens),\n        \"min\": int(lens.min()),\n        \"p50\": float(lens.median()),\n        \"p90\": float(lens.quantile(0.9)),\n        \"max\": int(lens.max()),\n    })\n\nif train is not None and SEQ_COL in train.columns:\n    print(\"TRAIN length stats\")\n    display(len_stats(train, SEQ_COL))\nif test is not None and SEQ_COL in test.columns:\n    print(\"TEST length stats\")\n    display(len_stats(test, SEQ_COL))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:19:16.491916Z","iopub.execute_input":"2026-02-02T15:19:16.492283Z","iopub.status.idle":"2026-02-02T15:19:16.500328Z","shell.execute_reply.started":"2026-02-02T15:19:16.492254Z","shell.execute_reply":"2026-02-02T15:19:16.498901Z"}},"outputs":[],"execution_count":null},{"id":"d982d3c0-5c47-46ed-89f7-89a0a55dc567","cell_type":"markdown","source":"## 5) Baseline predictor (DUMMY)\n\nReplace `predict_coords()` with your real model.\n\n**Contract**\n- input: RNA sequence string (length L)\n- output: `np.ndarray` shape `(L, 3)` float32\n","metadata":{}},{"id":"dd326222-b80a-4b6b-9024-6f76184e2a52","cell_type":"code","source":"import numpy as np\n\ndef predict_coords(sequence: str) -> np.ndarray:\n    # Dummy baseline: returns all-zeros coordinates with correct shape.\n    # Replace this with your real 3D model inference.\n    seq = str(sequence)\n    L = len(seq)\n    xyz = np.zeros((L, 3), dtype=np.float32)\n    return xyz\n\n# sanity check\nif test is not None and SEQ_COL in test.columns:\n    for i in range(min(3, len(test))):\n        seq = test.loc[i, SEQ_COL]\n        xyz = predict_coords(seq)\n        print(i, \"L=\", len(seq), \"xyz shape=\", xyz.shape, \"dtype=\", xyz.dtype)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:19:26.341439Z","iopub.execute_input":"2026-02-02T15:19:26.342212Z","iopub.status.idle":"2026-02-02T15:19:26.349236Z","shell.execute_reply.started":"2026-02-02T15:19:26.342178Z","shell.execute_reply":"2026-02-02T15:19:26.348039Z"}},"outputs":[],"execution_count":null},{"id":"1be37437-1a03-4cdc-ad80-0b961fd748b5","cell_type":"markdown","source":"## 6) Build submission\n\nSubmission formats vary. This template supports two common cases:\n\n**Case A — one row per ID** with a string payload (e.g. JSON-encoded coords).  \n**Case B — multiple rows per ID**, one per nucleotide with x/y/z columns.\n\nWe detect Case B if `x,y,z` (or `X,Y,Z`) exist in `sample_submission.csv`.\nIf it fails, edit the builder logic in this cell.\n","metadata":{}},{"id":"625fb8d6-57bb-45f1-a2f7-b272a7dc62a2","cell_type":"code","source":"import json\nimport numpy as np\nimport pandas as pd\n\nassert sub is not None, \"sample_submission.csv not found. Set DATA_DIR correctly.\"\n\ndef guess_id_col(df):\n    for c in [\"id\", \"target_id\", \"sequence_id\", \"rna_id\"]:\n        if c in df.columns:\n            return c\n    return df.columns[0]\n\n# Align ID/SEQ cols with test if needed\nif test is not None:\n    if ID_COL not in test.columns:\n        ID_COL = guess_id_col(test)\n    if SEQ_COL not in test.columns:\n        for c in [\"sequence\", \"seq\", \"rna_sequence\"]:\n            if c in test.columns:\n                SEQ_COL = c\n                break\n\nprint(\"Using ID_COL:\", ID_COL)\nprint(\"Using SEQ_COL:\", SEQ_COL)\n\ndisplay(sub.head(10))\nprint(\"sub columns:\", sub.columns.tolist())\n\nsub_cols = sub.columns.tolist()\ncase_b = all(c in sub_cols for c in [\"x\", \"y\", \"z\"]) or all(c in sub_cols for c in [\"X\", \"Y\", \"Z\"])\n\nrows = []\n\nif case_b:\n    xcol = \"x\" if \"x\" in sub_cols else \"X\"\n    ycol = \"y\" if \"y\" in sub_cols else \"Y\"\n    zcol = \"z\" if \"z\" in sub_cols else \"Z\"\n    idx_col = None\n    for c in [\"residue_index\", \"index\", \"i\", \"pos\", \"position\"]:\n        if c in sub_cols:\n            idx_col = c\n            break\n    if idx_col is None:\n        idx_col = \"residue_index\"  # fallback\n\n    for _, r in test.iterrows():\n        rid = r[ID_COL]\n        seq = r[SEQ_COL]\n        xyz = predict_coords(seq)\n        for i in range(xyz.shape[0]):\n            rows.append({\n                ID_COL: rid,\n                idx_col: i,\n                xcol: float(xyz[i,0]),\n                ycol: float(xyz[i,1]),\n                zcol: float(xyz[i,2]),\n            })\n\n    submission = pd.DataFrame(rows)\n    # Ensure required cols exist and order matches sample submission\n    for c in sub_cols:\n        if c not in submission.columns:\n            submission[c] = np.nan\n    submission = submission[sub_cols]\n\nelse:\n    id_col_sub = guess_id_col(sub)\n    target_cols = [c for c in sub_cols if c != id_col_sub]\n    assert len(target_cols) >= 1, \"Could not infer target column(s) from sample_submission.\"\n    target_col = target_cols[0]\n    print(\"Case A detected. Using target column:\", target_col)\n\n    for _, r in test.iterrows():\n        rid = r[ID_COL]\n        seq = r[SEQ_COL]\n        xyz = predict_coords(seq)\n        payload = json.dumps(xyz.tolist())  # [[x,y,z], ...]\n        rows.append({id_col_sub: rid, target_col: payload})\n\n    submission = pd.DataFrame(rows)\n\nsubmission.to_csv(\"submission.csv\", index=False)\nprint(\"Saved submission.csv with shape:\", submission.shape)\ndisplay(submission.head(10))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T15:19:29.576063Z","iopub.execute_input":"2026-02-02T15:19:29.576379Z","iopub.status.idle":"2026-02-02T15:19:29.608202Z","shell.execute_reply.started":"2026-02-02T15:19:29.576336Z","shell.execute_reply":"2026-02-02T15:19:29.606786Z"}},"outputs":[],"execution_count":null},{"id":"de0a5b64-91fb-4481-847b-6b82f6fdaf53","cell_type":"markdown","source":"## 7) Next steps\n\n1. Replace dummy inference with your model.\n2. Add CV + local TM-score using the official metric code (host refers to the metric notebook).\n3. Add strict leakage safeguards (release-date cutoff) and log external sources.\n","metadata":{}}]}