{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118765,"databundleVersionId":15231210,"sourceType":"competition"}],"dockerImageVersionId":31239,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Dataset Enrichment via GraphQL (RCSB PDB API)\n\nThis notebook augments sequence-level training data with structure-aware, evolutionary, and literature-linked metadata retrieved from the **RCSB Protein Data Bank GraphQL API**. Using GraphQL enables precise, schema-driven queries and avoids downloading large, unnecessary metadata tables. The resulting enriched dataset is provided in the notebook’s output directory as a preprocessed CSV, ready for downstream use.\n\n**What metadata is added**\n- **Experimental method** (e.g. **SOLUTION NMR**, **X-ray diffraction**, **Cryo-EM**)\n- **Primary citation information** (**PubMed ID**, **DOI**)\n- **Source organism taxonomy** (NCBI taxonomy ID and scientific name)\n- **Polymer source and evolutionary context** from PDB entity records\n\nAll queried fields are merged into the training dataframe using `rcsb_id` as the key. When metadata is unavailable for a given entry (e.g. missing citations or organism annotations), values are explicitly filled with **NaNs** to preserve dataframe shape and ensure compatibility with downstream filtering, modeling, and batching workflows.\n\n---\n\n### Why these features are useful\n\n**Experimental method (for TBM and template selection)**\n- In **template-based modeling (TBM)**, not all structural templates are equally informative.\n- Experimental method provides a coarse but useful proxy for structural reliability and representational bias:\n  - **X-ray / Cryo-EM** structures often provide well-defined global folds suitable for rigid-template alignment.\n  - **NMR** structures may represent conformational ensembles that are valuable for flexibility-aware modeling but less optimal as single static templates.\n- This metadata enables:\n  - Filtering MMseqs2 hits to prioritize templates solved by specific methods.\n  - Weighting candidate templates differently during TBM rather than treating all hits as equivalent.\n  - Method-aware benchmarking to avoid conflating modeling performance with experimental artifacts.\n\n**Evolutionary and organism-level context (for MMseqs2 filtering and hit prioritization)**\n- Taxonomic annotations provide evolutionary signal that complements sequence similarity scores from **MMseqs2**.\n- Closely related organisms often yield templates with more transferable structural and functional features, even at similar sequence identity.\n- Evolutionary metadata enables:\n  - Filtering or re-ranking MMseqs2 hits based on phylogenetic proximity rather than raw similarity alone.\n  - Emphasizing representative templates from specific clades while down-weighting redundant or evolutionarily distant matches.\n  - More principled template selection in low-identity regimes where sequence similarity alone is ambiguous.\n\n**DOI and PubMed identifiers (for RAG and interpretability)**\n- Citation metadata creates a direct bridge between structured structural data and the primary literature.\n- **PubMed IDs** and **DOIs** can be used in **retrieval-augmented generation (RAG) pipelines** to fetch abstracts or full-text articles associated with selected templates.\n- This enables automated retrieval of experimental details, functional annotations, and known limitations to support literature-grounded interpretation of modeling decisions.\n\n---\n\n### New dataset: train_seqs_extended_pdbdata.csv\n\nFor each entry in `train_seqs[\"target_id\"]`, the corresponding PDB ID is parsed and used to issue a targeted GraphQL query against the PDB `entries` endpoint. Because not all `target_id` values correspond to valid PDB entries (or may be missing entirely), the pipeline explicitly checks for invalid or missing identifiers and fills all unavailable fields with NaNs.\n\nThe resulting columns—`experimental_method`, `pubmed_id`, `doi`, `ncbi_taxonomy_id`, `organism_name`, and `bound_components`—are assigned to the dataframe in a single pass after querying, avoiding per-row overwrites and enabling clean downstream use.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import time\n\nimport pandas as pd\nimport numpy as np\n\nimport random\n#from Bio import pairwise2\n#from Bio.Seq import Seq\n%pip install rcsb-api\n\nfrom tqdm import tqdm\nfrom rcsbapi.data import DataQuery as Query\nfrom rcsbapi.data import DataQuery as Query\nimport json  # for easy-to-read output\n\nfrom scipy.spatial.transform import Rotation as R\nfrom sklearn.preprocessing import normalize\nfrom scipy. spatial import distance_matrix\nimport warnings\nwarnings.filterwarnings('ignore')\n\nprint(\"\\nLoading data files...\")\ntrain_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/train_sequences.csv')\nvalid_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/validation_sequences.csv')\ntest_seqs = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv')\ntrain_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/train_labels.csv')\nvalid_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/validation_labels.csv')\n\nprint(f\"Loaded {len(train_seqs)} training sequences, {len(valid_seqs)} validation sequences, and {len(test_seqs)} test sequences\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T20:16:34.330105Z","iopub.execute_input":"2026-01-12T20:16:34.330534Z","iopub.status.idle":"2026-01-12T20:16:48.75038Z","shell.execute_reply.started":"2026-01-12T20:16:34.330503Z","shell.execute_reply":"2026-01-12T20:16:48.749034Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#old train_seqs\ntrain_seqs","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T20:16:48.752324Z","iopub.execute_input":"2026-01-12T20:16:48.752692Z","iopub.status.idle":"2026-01-12T20:16:48.784658Z","shell.execute_reply.started":"2026-01-12T20:16:48.752659Z","shell.execute_reply":"2026-01-12T20:16:48.783741Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"experimental_methods = []\npubmed_ids = []\ndois = []\nncbi_taxonomy_ids = []\norganism_names = []\n\nfor pdb_id in tqdm(train_seqs[\"target_id\"], desc=\"Annotating PDB entries\"):\n    \n    # Skip missing/invalid IDs\n    if not isinstance(pdb_id, str) or not pdb_id.strip():\n        experimental_methods.append(np.nan)\n        pubmed_ids.append(np.nan)\n        dois.append(np.nan)\n        ncbi_taxonomy_ids.append(np.nan)\n        organism_names.append(np.nan)\n        continue\n    \n    pdb_id = pdb_id.strip().upper()\n    \n    query = Query(\n        input_type=\"entries\",\n        input_ids=[pdb_id],\n        return_data_list=[\n            \"exptl.method\",\n            \"rcsb_primary_citation.pdbx_database_id_PubMed\",\n            \"rcsb_primary_citation.pdbx_database_id_DOI\",\n            \"polymer_entities.rcsb_entity_source_organism.ncbi_taxonomy_id\",\n            \"polymer_entities.rcsb_entity_source_organism.ncbi_scientific_name\",\n        ],\n    )\n    \n    return_data = query.exec()\n    entries = (return_data.get(\"data\") or {}).get(\"entries\") or []\n    \n    # If the API returns nothing for this ID, fill NaNs and move on\n    if len(entries) == 0:\n        experimental_methods.append(np.nan)\n        pubmed_ids.append(np.nan)\n        dois.append(np.nan)\n        ncbi_taxonomy_ids.append(np.nan)\n        organism_names.append(np.nan)\n        continue\n    \n    entry = entries[0]\n    \n    exptl0 = (entry.get(\"exptl\") or [{}])[0]\n    experimental_methods.append(exptl0.get(\"method\", np.nan))\n    \n    citation = entry.get(\"rcsb_primary_citation\") or {}\n    pubmed_ids.append(citation.get(\"pdbx_database_id_PubMed\", np.nan))\n    dois.append(citation.get(\"pdbx_database_id_DOI\", np.nan))\n    \n    poly0 = (entry.get(\"polymer_entities\") or [{}])[0]\n    org0  = (poly0.get(\"rcsb_entity_source_organism\") or [{}])[0]\n    ncbi_taxonomy_ids.append(org0.get(\"ncbi_taxonomy_id\", np.nan))\n    organism_names.append(org0.get(\"ncbi_scientific_name\", np.nan))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T22:43:18.240712Z","iopub.execute_input":"2026-01-12T22:43:18.241066Z","iopub.status.idle":"2026-01-12T23:31:03.197857Z","shell.execute_reply.started":"2026-01-12T22:43:18.241039Z","shell.execute_reply":"2026-01-12T23:31:03.196917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs[\"experimental_method\"] = experimental_methods\ntrain_seqs[\"pubmed_id\"] = pubmed_ids\ntrain_seqs[\"doi\"] = dois\ntrain_seqs[\"ncbi_taxonomy_id\"] = ncbi_taxonomy_ids\ntrain_seqs[\"organism_name\"] = organism_names","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T23:45:13.232238Z","iopub.execute_input":"2026-01-12T23:45:13.233022Z","iopub.status.idle":"2026-01-12T23:45:13.242366Z","shell.execute_reply.started":"2026-01-12T23:45:13.232994Z","shell.execute_reply":"2026-01-12T23:45:13.241359Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#new train_seqs\ntrain_seqs","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T23:45:14.696029Z","iopub.execute_input":"2026-01-12T23:45:14.696452Z","iopub.status.idle":"2026-01-12T23:45:14.7176Z","shell.execute_reply.started":"2026-01-12T23:45:14.696379Z","shell.execute_reply":"2026-01-12T23:45:14.716356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T23:45:31.40878Z","iopub.execute_input":"2026-01-12T23:45:31.4091Z","iopub.status.idle":"2026-01-12T23:45:31.424951Z","shell.execute_reply.started":"2026-01-12T23:45:31.409076Z","shell.execute_reply":"2026-01-12T23:45:31.423781Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#example query using GraphQL\nquery = Query(\n    input_type=\"entries\",\n    input_ids=[\"8Z1F\"],\n    return_data_list=[\n        \"exptl.method\",\n        \"rcsb_ma_qa_metric_global.ma_qa_metric_global.type\",\n        \"rcsb_ma_qa_metric_global.ma_qa_metric_global.value\",\n        \"rcsb_primary_citation.pdbx_database_id_PubMed\",\n        \"rcsb_primary_citation.pdbx_database_id_DOI\",\n        \"rcsb_entry_info.nonpolymer_bound_components\",\n        \"polymer_entities.rcsb_entity_source_organism.ncbi_taxonomy_id\",\n        \"polymer_entities.rcsb_entity_source_organism.ncbi_scientific_name\",\n    ]\n)\nreturn_data = query.exec()\nprint(return_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T23:46:18.746632Z","iopub.execute_input":"2026-01-12T23:46:18.746974Z","iopub.status.idle":"2026-01-12T23:46:19.31647Z","shell.execute_reply.started":"2026-01-12T23:46:18.746939Z","shell.execute_reply":"2026-01-12T23:46:19.315248Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_seqs.to_csv('/kaggle/working/trainseqs_extended_pdbdata.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-12T23:46:22.183234Z","iopub.execute_input":"2026-01-12T23:46:22.183596Z","iopub.status.idle":"2026-01-12T23:46:23.372142Z","shell.execute_reply.started":"2026-01-12T23:46:22.183571Z","shell.execute_reply":"2026-01-12T23:46:23.371438Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}