{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# display the passes\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-04-27T10:59:50.490016Z","iopub.execute_input":"2023-04-27T10:59:50.490349Z","iopub.status.idle":"2023-04-27T10:59:51.045607Z","shell.execute_reply.started":"2023-04-27T10:59:50.49032Z","shell.execute_reply":"2023-04-27T10:59:51.044477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# IA.txt\nInformation Accretion for each term. This is used to weight precision and recall (see Evaluation)\n\nGO_term is like a unique number for a \"feature\". For example, the feature \"related to chloroplasts\" is numbered GO:0009507, and the feature \"related to membrane transport\" is numbered GO:0070633\n\nGO_termは「特徴」の固有番号である。例を挙げると「葉緑体に関係している」という特徴にはGO:0009507という番号が与えられ、「膜輸送に関係している」という特徴にはGO:0070633という番号が与えられている。","metadata":{}},{"cell_type":"code","source":"path_of_IA_file = '/kaggle/input/cafa-5-protein-function-prediction/IA.txt'\ndf_IA = pd.read_csv(path_of_IA_file , sep = '\\t', header = None).set_axis([\"GO_term\", \"weight\"], axis=1)\ndisplay(df_IA.head())\nsum_of_weight = (df_IA.iloc[:,1] == 0).sum()\nprint(sum_of_weight, 'count zero weight categories ', '%.1f - percent '%( sum_of_weight / len(df_IA)*100 ) )","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:11:36.908484Z","iopub.execute_input":"2023-04-27T11:11:36.908934Z","iopub.status.idle":"2023-04-27T11:11:36.943353Z","shell.execute_reply.started":"2023-04-27T11:11:36.908894Z","shell.execute_reply":"2023-04-27T11:11:36.942265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# sample_submission.tsv\n\nAccession_ID is the \"unique number of a protein\" as defined in Uniport, a comprehensive database of proteins. Specifically, the Accession ID of P-glycoprotein (a protein that transports membrane proteins) is P08183.\n\nAccession_IDはUniportというタンパク質の網羅的データベースにおいて定義されている「タンパク質の固有番号」である。具体的に、P-glycoprotein（膜輸送をおこなうタンパク質）のAccession IDはP08183である。","metadata":{}},{"cell_type":"markdown","source":"This competition quantifies how much a protein (Accession_ID) has a certain feature (GO_term).\n\n本コンペでは、あるタンパク質（Accession_ID）がある特徴（GO_term）をどれほど有しているかを定量する。","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv('../input/cafa-5-protein-function-prediction/sample_submission.tsv', sep='\\t',\n                  header=None, error_bad_lines=False).set_axis([\"Accession_ID\", \"GO_term\",\"prediction\"], axis=1)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:24:04.703956Z","iopub.execute_input":"2023-04-27T11:24:04.704323Z","iopub.status.idle":"2023-04-27T11:24:04.954661Z","shell.execute_reply.started":"2023-04-27T11:24:04.704292Z","shell.execute_reply":"2023-04-27T11:24:04.95374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train\n## train_taxonomy.tsv\n\nEntryID, like Accession_ID, indicates a \"unique protein number\"; there appear to be two types, 6-digit and 10-digit, but they are almost same.\n\nEntryIDはAccession_IDと同様に「タンパク質の固有番号」を示している。6桁と10桁の2種があるようだが、違いは特にないようだ","metadata":{}},{"cell_type":"markdown","source":"The TaxonomyID represents the \"unique number of the species\".　9606 indicates human and 559292 indicates budding yeast.\n\nTaxonomyIDは「種の固有番号」を表している。　9606はヒトを示し、559292は出芽酵母を示している。","metadata":{}},{"cell_type":"code","source":"train_tax = pd.read_csv('../input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv', sep='\\t', error_bad_lines=False)\ntrain_tax.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-26T13:32:13.141756Z","iopub.execute_input":"2023-04-26T13:32:13.142195Z","iopub.status.idle":"2023-04-26T13:32:13.23921Z","shell.execute_reply.started":"2023-04-26T13:32:13.142155Z","shell.execute_reply":"2023-04-26T13:32:13.238126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## train_terms.tsv\n\naspect indicates which of the three major categories it belongs to: biological process (BPO), cellular component (CCO), and molecular function (MFO).\n\naspectは、biological process（生物学的プロセス：BPO）、cellular component（細胞の構成要素：CCO）、molecular function（分子機能：MFO）という大きな三つの括りのうちのどれに属するかを示している。","metadata":{}},{"cell_type":"code","source":"train_terms = pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv', sep='\\t', error_bad_lines=False)\ntrain_terms.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:41:54.485393Z","iopub.execute_input":"2023-04-27T11:41:54.485727Z","iopub.status.idle":"2023-04-27T11:41:57.737782Z","shell.execute_reply.started":"2023-04-27T11:41:54.485701Z","shell.execute_reply":"2023-04-27T11:41:57.73602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# only three aspects\ntrain_terms[\"aspect\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-04-27T11:42:43.0049Z","iopub.execute_input":"2023-04-27T11:42:43.005469Z","iopub.status.idle":"2023-04-27T11:42:43.290027Z","shell.execute_reply.started":"2023-04-27T11:42:43.005437Z","shell.execute_reply":"2023-04-27T11:42:43.289237Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## train_sequences.fasta\n\nThe FASTA format is a text-based format for representing either nucleotide or amino acid sequences, where nucleotides or amino acids are represented using a single letter code. In this data set, the amino acid sequence for each protein is shown; P20536 is the AccessionID for the protein Uracil-DNA glycosylase, which is composed of 218 amino acids in a row. Counting the number of sequences (MNS~FIY) shown below, we can confirm that the number is indeed 218.\n\nFASTA形式は、ヌクレオチド配列またはアミノ酸配列のいずれかを表すためのテキストベースの形式であり、ヌクレオチドまたはアミノ酸は1文字のコードを使用して表される。本データでは、各タンパク質のアミノ酸配列を示している。P20536はUracil-DNA glycosylaseというタンパク質のAccessionIDであり、このタンパク質はアミノ酸が218個並んで構成されている。下に示した配列（MNS~FIY）の個数を数えると確かに218個になっていることが確認できる。","metadata":{}},{"cell_type":"code","source":"!head '/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta'","metadata":{"execution":{"iopub.status.busy":"2023-04-27T12:00:19.032083Z","iopub.execute_input":"2023-04-27T12:00:19.032488Z","iopub.status.idle":"2023-04-27T12:00:19.359036Z","shell.execute_reply.started":"2023-04-27T12:00:19.032453Z","shell.execute_reply":"2023-04-27T12:00:19.357786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## go-basic.obo\n\nGO term details (attributes).\n\nGO termの詳細（属性）を示している。","metadata":{}},{"cell_type":"code","source":"!head '/kaggle/input/cafa-5-protein-function-prediction/Train/go-basic.obo' -n 50","metadata":{"execution":{"iopub.status.busy":"2023-04-27T12:17:05.01411Z","iopub.execute_input":"2023-04-27T12:17:05.014484Z","iopub.status.idle":"2023-04-27T12:17:05.302299Z","shell.execute_reply.started":"2023-04-27T12:17:05.014447Z","shell.execute_reply":"2023-04-27T12:17:05.300602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test\n## testsuperset-taxon-list.tsv\n\nThe ID represents the TaxonomyID, or \"unique number of the species. Species on the right side is the specific name of the organism.\n\nIDはTaxonomyID、つまり「種の固有番号」を表している。右側のSpeciesは具体的に生物名が記されている。","metadata":{}},{"cell_type":"code","source":"test_superset_taxon = pd.read_csv('../input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset-taxon-list.tsv', sep='\\t', error_bad_lines=False, encoding= 'unicode_escape')\ntest_superset_taxon.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-26T13:39:02.546923Z","iopub.execute_input":"2023-04-26T13:39:02.547322Z","iopub.status.idle":"2023-04-26T13:39:02.564153Z","shell.execute_reply.started":"2023-04-26T13:39:02.547288Z","shell.execute_reply":"2023-04-26T13:39:02.562778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## testsuperset.fasta\n\nSame as train_sequences.fasta. The second number indicates the \"unique number of the species. (10090 is a house mouse)\n\ntrain_sequences.fastaと同様。二つ目の数字は「種の固有番号」を示す。（10090はハツカネズミ）","metadata":{}},{"cell_type":"code","source":"!head \"/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset.fasta\"","metadata":{"execution":{"iopub.status.busy":"2023-04-27T12:10:37.261162Z","iopub.execute_input":"2023-04-27T12:10:37.261565Z","iopub.status.idle":"2023-04-27T12:10:37.546393Z","shell.execute_reply.started":"2023-04-27T12:10:37.261531Z","shell.execute_reply":"2023-04-27T12:10:37.545544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}