{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118765,"databundleVersionId":15231210,"sourceType":"competition"}],"dockerImageVersionId":31234,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on January 08, 2025. By Prata, Marília (mpwolke)","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:32:39.883952Z","iopub.execute_input":"2026-01-09T15:32:39.884627Z","iopub.status.idle":"2026-01-09T15:32:58.829229Z","shell.execute_reply.started":"2026-01-09T15:32:39.884593Z","shell.execute_reply":"2026-01-09T15:32:58.828184Z"},"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Competititon Citation\n\n@misc{stanford-rna-3d-folding-2,\n\n    author = {Przemek Porebski  and Rhiju Das and Walter Reade and Ashley Oldacre},\n    \n    title = {Stanford RNA 3D Folding Part 2},\n    year = {2026},\n    howpublished = {\\url{https://kaggle.com/competitions/stanford-rna-3d-folding-2}},\n    \n    note = {Kaggle}\n}","metadata":{}},{"cell_type":"markdown","source":"## Train (train_sequences) file\n\ntarget_id - (string) An arbitrary identifier. In train_sequences.csv, this is formatted as pdb_id_chain_id, where pdb_id is the id of the entry in the Protein Data Bank and chain_id is the chain id of the monomer in the pdb file.\n\nsequence - (string) The RNA sequence of all chains in the target, concatenated together according to stoichiometry\n\ntemporal_cutoff - (string) The date in yyyy-mm-dd format that the sequence was or will be published.\ndescription - (string) Details of the origins of the sequence. For PDB entries, this is the entry title.\n\n**stoichiometry** - (string) the chains used for the target. These take the form of {chain:number}, where chain corresponds to the author-defined chain in all_sequences, joined with a semicolon delimiter (;).\n\nall_sequences - (string) FASTA-formatted sequences of all molecular chains present in the experimentally solved structure. May include multiple copies of the target RNA (look for the word \"Chains\" in the header) and/or partners like other RNAs or proteins or DNA. You **don't need to make predictions for all these molecules**, just the ones **specified in stoichiometry** which are concatenated in sequence. Can be parsed into a dictionary with extra/parse_fasta_py.py.\n\nligand_ids - (string) three-letter names in PDB chemical component dictionary of any small molecule ligands solved in the experimental structure, joined with a semicolon delimiter (;). You don't need to make predictions for these molecules.\n\nligand_SMILES - (string) SMILES strings giving chemical structures of any small molecule ligands solved in the experimental structure, joined with a semicolon delimiter (;).\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/train_sequences.csv')\ntrain.tail(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:34:18.289064Z","iopub.execute_input":"2026-01-09T15:34:18.289565Z","iopub.status.idle":"2026-01-09T15:34:19.011969Z","shell.execute_reply.started":"2026-01-09T15:34:18.289536Z","shell.execute_reply":"2026-01-09T15:34:19.011286Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train['description'].value_counts()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:34:25.717542Z","iopub.execute_input":"2026-01-09T15:34:25.71789Z","iopub.status.idle":"2026-01-09T15:34:25.737132Z","shell.execute_reply.started":"2026-01-09T15:34:25.717861Z","shell.execute_reply":"2026-01-09T15:34:25.73621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#!pip install pubchempy","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:10:51.768569Z","iopub.execute_input":"2026-01-09T15:10:51.769926Z","iopub.status.idle":"2026-01-09T15:10:56.891206Z","shell.execute_reply.started":"2026-01-09T15:10:51.769884Z","shell.execute_reply":"2026-01-09T15:10:56.889035Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Stoichiometry\n\n\"Stoichiometry is based on the law of conservation of mass; the total mass of reactants must equal the total mass of products, so the relationship between reactants and products must form a ratio of positive integers. This means that if the amounts of the separate reactants are known, then the amount of the product can be calculated.\"\n\nhttps://en.wikipedia.org/wiki/Stoichiometry","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/test_sequences.csv')\ntest.tail(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:34:58.380512Z","iopub.execute_input":"2026-01-09T15:34:58.380822Z","iopub.status.idle":"2026-01-09T15:34:58.398675Z","shell.execute_reply.started":"2026-01-09T15:34:58.380798Z","shell.execute_reply":"2026-01-09T15:34:58.397572Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Validation sequences file\n\nThe validation_sequences.csv (which is the same as test_sequences.csv) comprise targets released after May 29, 2025, the final submission date of the last Stanford RNA 3D Folding competition, and up to December 17, 2025. These were further filtered to have composition of at least 40% RNA and unique sequences (up to sequence identity 90%).\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data","metadata":{}},{"cell_type":"code","source":"val = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/validation_sequences.csv')\nval.tail(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:35:03.873386Z","iopub.execute_input":"2026-01-09T15:35:03.873742Z","iopub.status.idle":"2026-01-09T15:35:03.891822Z","shell.execute_reply.started":"2026-01-09T15:35:03.8737Z","shell.execute_reply":"2026-01-09T15:35:03.890986Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Submission file\n\nSame format as train_labels.csv but with five sets of coordinates for each of your five predicted structures (x_1,y_1,z_1,x_2,y_2,z_2,…x_5,y_5,z_5).\n\nYou must **submit five sets of coordinates**.\n\nNote that x,y,z are clipped between -999.999 and 9999.999 before scoring, due to use of a legacy PDB format that has maximal 8 characters for coordinates.\n\nchain and copy do not have to be provided.\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/sample_submission.csv')\nsub.tail()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:35:09.889765Z","iopub.execute_input":"2026-01-09T15:35:09.890741Z","iopub.status.idle":"2026-01-09T15:35:09.927742Z","shell.execute_reply.started":"2026-01-09T15:35:09.890683Z","shell.execute_reply":"2026-01-09T15:35:09.927037Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"![](https://www.medschoolcoach.com/wp-content/uploads/2022/11/TheDNADoubleHelix-Figure1.jpg)","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/train_labels.csv')\ntrain_labels.tail(3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:35:17.308896Z","iopub.execute_input":"2026-01-09T15:35:17.30962Z","iopub.status.idle":"2026-01-09T15:35:27.56007Z","shell.execute_reply.started":"2026-01-09T15:35:17.309586Z","shell.execute_reply":"2026-01-09T15:35:27.559323Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import plotly.express as px\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\n\n#Two lines Required to Plot Plotly\nimport plotly.io as pio\npio.renderers.default = 'iframe'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:02:11.982087Z","iopub.execute_input":"2026-01-09T14:02:11.982936Z","iopub.status.idle":"2026-01-09T14:02:15.189272Z","shell.execute_reply.started":"2026-01-09T14:02:11.982907Z","shell.execute_reply":"2026-01-09T14:02:15.188153Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## validation_labels.csv - experimental structures.","metadata":{}},{"cell_type":"code","source":"val_labels = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/validation_labels.csv')\nval_labels.tail(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:36:31.438511Z","iopub.execute_input":"2026-01-09T14:36:31.440109Z","iopub.status.idle":"2026-01-09T14:36:31.654502Z","shell.execute_reply.started":"2026-01-09T14:36:31.440036Z","shell.execute_reply":"2026-01-09T14:36:31.653507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## resname - (character) The RNA nucleotide ( A, C, G, or U) for the residue.\n\n**Adenine, Cytosine, Guanine and Uracil**. \n\nBy the way, not a very good 3D chart. ","metadata":{}},{"cell_type":"markdown","source":"### STRUCTURE OF THE THERMUS THERMOPHILUS\n\nEven a Thermus thermophilus looks better than my 3D : (\n\n![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcSVeL5B0Jn2NrZkY9kHMYFWOLPIlZ-GbTnm9g&s)\n\nhttps://www.ncbi.nlm.nih.gov/Structure/pdb/1IBK","metadata":{}},{"cell_type":"code","source":"#Code by Anmorgul https://www.kaggle.com/anmorgul/strange-pattern-cottonwood-willow\n#https://www.kaggle.com/code/mpwolke/roosevelt-forest-of-northern-colorado-charts\n\nfor i in range(4,5):\n    fig = px.scatter_3d(val_labels, x='x_1', y='y_1', z='z_1',\n                  color='resname', size_max=8, width=800, height=600, opacity=0.9, template=\"plotly_dark\")\n    fig.update_layout(\n        font_size=8,\n        legend_font_size=16,)\n    fig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:02:33.80761Z","iopub.execute_input":"2026-01-09T14:02:33.807984Z","iopub.status.idle":"2026-01-09T14:02:36.29562Z","shell.execute_reply.started":"2026-01-09T14:02:33.807956Z","shell.execute_reply":"2026-01-09T14:02:36.294582Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## meta file - Canonical AGCU (Adenine, Guanine, Cytosine, and Uracil)\n\n\"The term **\"canonical AGCU\"** most likely refers to the four primary canonical RNA bases: Adenine, Guanine, Cytosine, and Uracil. These are the fundamental units of the genetic code in ribonucleic acid (RNA).\" \n\nThe additional file extra/rna_metadata.csv contains data extracted from all RNA and RNA/DNA hybrid structures up to December 17,2025. These metadata were used to filter the structures that were included in the {train,test,validation}_sequences.csv using the following **criteria**:\n\n**canonical ACGU** residues or modified residues that can be mapped to canonical using either PDB chemical component dictionary or NAKB mapping\n\nno undefined (N) residues and no T (for hybrid NA)\n\nno more than 25% of residues that were modified / non-canonical\n\nat least 50% residues reported in the sequence were modeled/observed","metadata":{}},{"cell_type":"code","source":"meta = pd.read_csv('/kaggle/input/stanford-rna-3d-folding-2/extra/rna_metadata.csv')\nmeta.tail(2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T16:04:33.589312Z","iopub.execute_input":"2026-01-09T16:04:33.589627Z","iopub.status.idle":"2026-01-09T16:04:35.634577Z","shell.execute_reply.started":"2026-01-09T16:04:33.589601Z","shell.execute_reply":"2026-01-09T16:04:35.633909Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## unmapped_canonical","metadata":{}},{"cell_type":"code","source":"#plotting the piechart for loan_status column.\n\ncanonical = meta['unmapped_canonical'].value_counts()\nplt.pie(canonical.values,\n        labels=canonical.index,\n        autopct='%1.1f%%')\nplt.title('Unmapped Canonical')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T16:09:15.129208Z","iopub.execute_input":"2026-01-09T16:09:15.129793Z","iopub.status.idle":"2026-01-09T16:09:15.266364Z","shell.execute_reply.started":"2026-01-09T16:09:15.129765Z","shell.execute_reply":"2026-01-09T16:09:15.265517Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Stacking as a Driver for RNA Structure\n\nCitation: Thoughts on how to think (and talk) about RNA structure\nApril 2022Proceedings of the National Academy of Sciences 119(17):e2112677119 - DOI:10.1073/pnas.2112677119\n\nAuthors: Quentin Vicens and Jeffrey S Kieft\n\n\n\"To understand RNA structure, **one must therefore understand base stacking**. What is the experimental evidence that base stacking drives RNA structure? Starting in the 1950s, biophysical studies of RNA oligonucleotides in solution revealed that unpaired bases stack in helical configurations. Even a single ApA dinucleotide adopts the “beginning of a single-strand helix”.\"\n\n\"This helix was modeled in 1975 on the basis of the crystal structure of ApApA. In fact, poly-adenosine [poly(A)] forms a parallel double helix, even though it cannot form Watson–Crick pairs. This “second double helix” was first reported by **Watson and Crick in 1961**, and its structure was confirmed many decades later.\"\n\n### Automatic Predictions of an RNA Secondary Structure: Approach with Caution!\n\n### Mfold - m refers to multiple\n\n\"The **‘mfold’ software for RNA folding** was developed in the late 1980s. The ‘m’ simply refers to ‘multiple’. The core algorithm predicts a minimum free energy, ΔG, as well as minimum free energies for foldings that must contain any particular base pair.\"\n\nhttps://pmc.ncbi.nlm.nih.gov/articles/PMC169194/#:~:text=The%20'mfold'%20software%20for%20RNA,described%20(23%2C24).\n\n\n\"When researchers encounter a novel RNA sequence, they commonly **“Mfold” it**. Within seconds, a convenient prediction emerges of the lowest predicted free-energy secondary structure of that RNA, computed by a **thermodynamics-based algorithm**. Unfortunately, often this output from **Mfold** (or RNAfold, Sfold, and other similar tools is not presented as a prediction, resulting from an algorithm that includes assumptions, but as a representation of the true structure. Alternative pairing possibilities with theoretically higher free energies are generally ignored and the top prediction may even be propagated in the literature without supporting evidence. The danger of only considering the lowest free-energy structure and showcasing it as “The” secondary structure of a particular RNA is that the community ends up taking an untested possibility at face value.\"\n\n\"Unfortunately, the outputs from folding programs often substitute for a thorough experimental secondary structure determination. RNA secondary structure prediction software built on rigorous measurements of base-pair stability in different nearest neighbor contexts are valuable tools for rapidly assessing potential **Watson–Crick pairing** and generating new ideas about RNA structure. However, they consider all nucleotides as equally likely to be involved in secondary structure elements, which often leads to erroneous assumptions about base pairs.\"\n\n\"Furthermore, because these algorithms tend to maximize the predicted number of **Watson–Crick pairs** (equated with the lowest predicted free energy), they are most accurate when an RNA has many such pairs. The problem is, not all RNAs do.\"\n\n\nhttps://www.pnas.org/doi/10.1073/pnas.2112677119","metadata":{}},{"cell_type":"markdown","source":"## Fasta file - One single fasta file\n\nThat's not fast.","metadata":{}},{"cell_type":"code","source":"#https://stackoverflow.com/questions/29805642/learning-to-parse-a-fasta-file-with-python\n\ndef read_fasta(fp):\n        name, seq = None, []\n        for line in fp:\n            line = line.rstrip()\n            if line.startswith(\">\"):\n                if name: yield (name, ''.join(seq))\n                name, seq = line, []\n            else:\n                seq.append(line)\n        if name: yield (name, ''.join(seq))\n\nwith open('../input/stanford-rna-3d-folding-2/MSA/1NJN.MSA.fasta') as fp:\n    for name, seq in read_fasta(fp):\n        print(name, seq)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:41:59.656448Z","iopub.execute_input":"2026-01-09T15:41:59.656966Z","iopub.status.idle":"2026-01-09T15:41:59.947289Z","shell.execute_reply.started":"2026-01-09T15:41:59.656938Z","shell.execute_reply":"2026-01-09T15:41:59.945908Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## parse_fasta_py.py\n\nhttps://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data","metadata":{}},{"cell_type":"code","source":"#https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data\n#https://www.google.com/search?q=%27Dict%27+is+not+defined&sca_esv=8e3f949f37700064&sxsrf=ANbL-n59EeuQXxVeeRVevhh0K7lMpqSTeg%3A1767970426091&ei=ehZhaaupBZiY5OUPxrXAgQg&ved=0ahUKEwirzuGJ2_6RAxUYDLkGHcYaMIAQ4dUDCBE&uact=5&oq=%27Dict%27+is+not+defined&gs_lp=Egxnd3Mtd2l6LXNlcnAiFSdEaWN0JyBpcyBub3QgZGVmaW5lZDIGEAAYFhgeMgYQABgWGB4yBhAAGBYYHjIGEAAYFhgeMgYQABgWGB4yBhAAGBYYHjIGEAAYFhgeMgYQABgWGB4yCBAAGBYYChgeMggQABgWGAoYHkioH1DJEViSF3ACeACQAQCYAegBoAHoAaoBAzItMbgBA8gBAPgBAfgBApgCA6ACgwKoAhDCAggQABiwAxjvBcICCxAAGLADGKIEGIkFwgIHECMYJxjqAsICDRAjGIAEGCcYigUY6gLCAhMQIxjwBRiABBgnGMkCGIoFGOoCwgINEC4YgAQYJxiKBRjqAsICEBAjGPAFGIAEGCcYigUY6gLCAhcQABiABBiRAhi0AhjnBhiKBRjqAtgBAZgDCvEFjEG4yP0e9mGIBgGQBgW6BgYIARABGAGSBwUyLjAuMaAHjQeyBwMyLTG4B-0BwgcFMi0yLjHIBxSACAA&sclient=gws-wiz-serp\n\nfrom typing import Dict, List, Tuple\n\n\ndef parse_fasta(fasta_content: str) -> Dict[str, Tuple[str, List[str]]]:\n    \"\"\"\n    Parse FASTA content into dictionary.\n\n    Args:\n        fasta_content: Multi-line FASTA string with format:\n        >1A1T_1|Chain A[auth B]|SL3 STEM-LOOP RNA|\n        or\n        >104D_1|Chains A[auth A], B[auth B]|DNA/RNA (...)|\n\n    Returns:\n        Dictionary mapping auth chain_id to (sequence, list_of_auth_chain_ids)\n        Example: {\"A\": (\"ACGT\", [\"A\", \"B\"]), \"C\": (\"UGCA\", [\"C\"])}\n        The key is the auth chain ID, and the list contains all auth chain IDs for this sequence\n    \"\"\"\n    result = {}\n    lines = fasta_content.strip().split(\"\\n\")\n\n    i = 0\n    while i < len(lines):\n        line = lines[i].strip()\n\n        if line.startswith(\">\"):\n            # Parse new format header: >104D_1|Chains A[auth A], B[auth B]|...| or >1A1T_1|Chain A[auth B]|...|\n            # Extract the chains part (between first | and second |)\n            parts = line.split(\"|\")\n            if len(parts) < 2:\n                print(\"Warning: Malformed FASTA header:\", line)\n                auth_chain_ids = []\n                chains_part = \"\"\n            else:\n                chains_part = parts[1].strip()\n\n                # Extract auth chain IDs from patterns like \"Chain A[auth B]\" or \"Chains A[auth A], B[auth B] or just \"Chain A\" or \"Chains A, B\"\n                auth_chain_ids = []\n                replaced_chains_part = re.sub(r\"^Chains? \", \"\", chains_part)\n                chains = replaced_chains_part.split(\",\")\n                for chain in chains:\n                    auth_match = re.search(r\"\\[auth ([^\\]]+)\\]\", chain)\n                    if auth_match:\n                        auth_chain_ids.append(auth_match.group(1).strip())\n                    else:\n                        c = chain.strip()\n                        if c:\n                            auth_chain_ids.append(c)\n\n            if not auth_chain_ids:\n                print(\"Warning: Empty chains part:\", chains_part)\n                primary_auth_chain = None\n            else:\n                # Use the first auth chain ID as the key\n                primary_auth_chain = auth_chain_ids[0]\n\n            # Read sequence (next lines until next header or end)\n            sequence = \"\"\n            while (i + 1) < len(lines) and lines[i + 1].startswith(\">\") is False:\n                sequence += lines[i + 1].strip()\n                i += 1\n            result[primary_auth_chain] = (sequence, auth_chain_ids)\n\n        i += 1\n\n    return result","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:47:05.144782Z","iopub.execute_input":"2026-01-09T15:47:05.145077Z","iopub.status.idle":"2026-01-09T15:47:05.153935Z","shell.execute_reply.started":"2026-01-09T15:47:05.145054Z","shell.execute_reply":"2026-01-09T15:47:05.15292Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### What's next?","metadata":{}},{"cell_type":"code","source":"print(parse_fasta)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T15:51:12.849194Z","iopub.execute_input":"2026-01-09T15:51:12.850085Z","iopub.status.idle":"2026-01-09T15:51:12.854181Z","shell.execute_reply.started":"2026-01-09T15:51:12.850045Z","shell.execute_reply":"2026-01-09T15:51:12.853308Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#This is the cheminformatics package that will do most of the heavy lifting for us\n!pip install rdkit","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:19:26.730753Z","iopub.execute_input":"2026-01-09T14:19:26.732118Z","iopub.status.idle":"2026-01-09T14:19:35.928033Z","shell.execute_reply.started":"2026-01-09T14:19:26.732028Z","shell.execute_reply":"2026-01-09T14:19:35.926913Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#By Chemdatafarmer  https://www.kaggle.com/code/chemdatafarmer/additional-seh-data/notebook\n#By Meer Atif https://www.kaggle.com/code/meeratif/smiles-open-problems\n\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw, AllChem\nfrom rdkit import RDLogger","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:19:53.41089Z","iopub.execute_input":"2026-01-09T14:19:53.412061Z","iopub.status.idle":"2026-01-09T14:19:53.743805Z","shell.execute_reply.started":"2026-01-09T14:19:53.412003Z","shell.execute_reply":"2026-01-09T14:19:53.742448Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#10th row, 2nd colund\n\ntrain.iloc[10,1]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:50:10.073914Z","iopub.execute_input":"2026-01-09T14:50:10.075167Z","iopub.status.idle":"2026-01-09T14:50:10.081976Z","shell.execute_reply.started":"2026-01-09T14:50:10.07512Z","shell.execute_reply":"2026-01-09T14:50:10.080788Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#!pip install pysmiles","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:50:20.533252Z","iopub.execute_input":"2026-01-09T14:50:20.534218Z","iopub.status.idle":"2026-01-09T14:50:25.736569Z","shell.execute_reply.started":"2026-01-09T14:50:20.534188Z","shell.execute_reply":"2026-01-09T14:50:25.734977Z"},"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## UGGGGUUGGGGUUGGGGUUGGGGU!\n\n### I tried pychem, rdkit, pubchempy, deepchem, Molecular Descriptor Calculator (Mordred). However, I couldn't plot a single mol.\n\nAlso the helper function that's on the extra/parse_fasta_py.py of this competition data.","metadata":{}},{"cell_type":"code","source":"from rdkit import Chem\n\nsmiles_string = \"UGGGGUUGGGGUUGGGGUUGGGGU\" # Example with leading space\n\n# Trim whitespace\ncleaned_smiles = smiles_string.strip()\n\nmol = Chem.MolFromSmiles(cleaned_smiles)\n\nif mol is not None:\n    print(\"Molecule successfully parsed.\")\n    # Proceed with molecule processing\nelse:\n    print(f\"Failed to parse SMILES: '{cleaned_smiles}'\")\n    # Investigate the content of cleaned_smiles further","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-09T14:55:22.799283Z","iopub.execute_input":"2026-01-09T14:55:22.800153Z","iopub.status.idle":"2026-01-09T14:55:22.806918Z","shell.execute_reply.started":"2026-01-09T14:55:22.800122Z","shell.execute_reply":"2026-01-09T14:55:22.805709Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## No mol, just SMILES Parse Error: Failed parsing SMILES .\n\n### I'm not smilling.\n\n\"**SMILES Parse Error**: check for mistakes around position 1\" indicates that the very first character of your input string is not a valid component of a SMILES (Simplified Molecular Input Line Entry System) string.\" ","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nAnmorgul https://www.kaggle.com/anmorgul/strange-pattern-cottonwood-willow\n\nparse_fasta_py.py https://www.kaggle.com/competitions/stanford-rna-3d-folding-2/data","metadata":{}}]}