{
  "id": 674359,
  "title": "ambiguity of training data's cutoff",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/674359",
  "author_name": "WJui",
  "post_date": "2026-02-20T01:32:56.188000",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>According to the dataset description,</p>\n<blockquote>\n  <p>The train_sequences.csv has not been filtered for sequence redundancy and contains PDB entries released before May 29, 2025 …</p>\n</blockquote>\n<p>However, this image shows that the latest cutoff date is 2025-12-17</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24942389%2F063b9119dacf0196a80ee6f5c1594a25%2F20260220093031.jpg?generation=1771551056907558&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3408505,
      "postDate": "2026-02-20T19:24:46.653Z",
      "content": "<p>Thank you for raising this issue. We prepared the train/validation split using clustering at a 30% identity level and a temporal cutoff for each cluster. Since the initial description of the data preparation step was unclear, we will update it shortly. The <code>temporal_cutoff</code>  in the <code>train_sequences.csv</code> and <code>validation_sequences.csv</code> are the cutoffs associated with individual entries. </p>\n<p>To create the split, we assigned each <code>pdb_id</code> a date based on the oldest <code>temporal_cutoff</code> within its MMseqs2 cluster (at 30% identity). For example, if two entries—PDB1 (2020-01-01) and PDB2 (2025-12-17) belong to the same cluster, both will be assigned to the training set because the cluster's cutoff becomes 2020-01-01. We chose this method to minimize the presence of structures in the validation set that are too similar to the training data and reduce the leakage due to presence of possible homologs. </p>\n<p>You can reproduce this procedure using the metadata in the extra folder:</p>\n<pre><code>import pandas as pd\n\n# Read the chain metadata \nmetadata = pd.read_csv('extra/rna_metadata.csv', parse_dates=True)\n\n# Assign the oldest temporal cutoff for each cluster to all members of the cluster\ncluster_temporal_cutoff = metadata.groupby('mmseqs_0.300')['temporal_cutoff'].min()\ncluster_temporal_cutoff.name = 'cluster_temporal_cutoff'\nmetadata_with_cluster_cutoff = metadata.merge(cluster_temporal_cutoff, on='mmseqs_0.300')\n# For each PDB ID, get the minimum cluster temporal cutoff among its chains\npdbid_split_cutoff = metadata_with_cluster_cutoff.groupby('pdb_id')['cluster_temporal_cutoff'].min()\n\n# Split the PDB IDs into training and validation sets based on the cluster temporal cutoff\n# NOTE: the sequences undergo further filtering, so we have supersets of PDB IDs here\ntrain_superset = set(pdbid_split_cutoff[pdbid_split_cutoff &lt; '2025-05-29' ].index)\nvalidation_superset = set(pdbid_split_cutoff[pdbid_split_cutoff &gt;= '2025-05-29'].index)\n\ntrain_sequences = pd.read_csv('train_sequences.csv')\nvalidation_sequences = pd.read_csv('validation_sequences.csv')\n\ntrain_set = set(train_sequences['target_id'])\nvalidation_set = set(validation_sequences['target_id'])\n\n# Validate\nassert validation_set.isdisjoint(train_superset)\nassert train_set.isdisjoint(validation_superset)\nassert train_set.issubset(train_superset)\nassert validation_set.issubset(validation_superset)\n</code></pre>\n<p>Hope that clarifies the issue.</p>",
      "rawMarkdown": "Thank you for raising this issue. We prepared the train/validation split using clustering at a 30% identity level and a temporal cutoff for each cluster. Since the initial description of the data preparation step was unclear, we will update it shortly. The `temporal_cutoff`  in the `train_sequences.csv` and `validation_sequences.csv` are the cutoffs associated with individual entries. \n\nTo create the split, we assigned each `pdb_id` a date based on the oldest `temporal_cutoff` within its MMseqs2 cluster (at 30% identity). For example, if two entries—PDB1 (2020-01-01) and PDB2 (2025-12-17) belong to the same cluster, both will be assigned to the training set because the cluster's cutoff becomes 2020-01-01. We chose this method to minimize the presence of structures in the validation set that are too similar to the training data and reduce the leakage due to presence of possible homologs. \n\nYou can reproduce this procedure using the metadata in the extra folder:\n\n```python\nimport pandas as pd\n\n# Read the chain metadata \nmetadata = pd.read_csv('extra/rna_metadata.csv', parse_dates=True)\n\n# Assign the oldest temporal cutoff for each cluster to all members of the cluster\ncluster_temporal_cutoff = metadata.groupby('mmseqs_0.300')['temporal_cutoff'].min()\ncluster_temporal_cutoff.name = 'cluster_temporal_cutoff'\nmetadata_with_cluster_cutoff = metadata.merge(cluster_temporal_cutoff, on='mmseqs_0.300')\n# For each PDB ID, get the minimum cluster temporal cutoff among its chains\npdbid_split_cutoff = metadata_with_cluster_cutoff.groupby('pdb_id')['cluster_temporal_cutoff'].min()\n\n# Split the PDB IDs into training and validation sets based on the cluster temporal cutoff\n# NOTE: the sequences undergo further filtering, so we have supersets of PDB IDs here\ntrain_superset = set(pdbid_split_cutoff[pdbid_split_cutoff < '2025-05-29' ].index)\nvalidation_superset = set(pdbid_split_cutoff[pdbid_split_cutoff >= '2025-05-29'].index)\n\ntrain_sequences = pd.read_csv('train_sequences.csv')\nvalidation_sequences = pd.read_csv('validation_sequences.csv')\n\ntrain_set = set(train_sequences['target_id'])\nvalidation_set = set(validation_sequences['target_id'])\n\n# Validate\nassert validation_set.isdisjoint(train_superset)\nassert train_set.isdisjoint(validation_superset)\nassert train_set.issubset(train_superset)\nassert validation_set.issubset(validation_superset)\n```\nHope that clarifies the issue.\n",
      "votes": 3,
      "replies": [
        {
          "id": 3408605,
          "postDate": "2026-02-21T00:27:51.767Z",
          "content": "<p>Thank you for your clarification!</p>",
          "rawMarkdown": "Thank you for your clarification!"
        }
      ]
    },
    {
      "id": 3408192,
      "postDate": "2026-02-20T01:32:56.190Z",
      "content": "<p>According to the dataset description,</p>\n<blockquote>\n  <p>The train_sequences.csv has not been filtered for sequence redundancy and contains PDB entries released before May 29, 2025 …</p>\n</blockquote>\n<p>However, this image shows that the latest cutoff date is 2025-12-17</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24942389%2F063b9119dacf0196a80ee6f5c1594a25%2F20260220093031.jpg?generation=1771551056907558&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "According to the dataset description,\n>The train_sequences.csv has not been filtered for sequence redundancy and contains PDB entries released before May 29, 2025 ...\n\nHowever, this image shows that the latest cutoff date is 2025-12-17\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24942389%2F063b9119dacf0196a80ee6f5c1594a25%2F20260220093031.jpg?generation=1771551056907558&alt=media)",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3408505,
      "author_name": "Przemek Porebski",
      "author_url": "",
      "post_date": "2026-02-20T19:24:46.653000",
      "content": "<p>Thank you for raising this issue. We prepared the train/validation split using clustering at a 30% identity level and a temporal cutoff for each cluster. Since the initial description of the data preparation step was unclear, we will update it shortly. The <code>temporal_cutoff</code>  in the <code>train_sequences.csv</code> and <code>validation_sequences.csv</code> are the cutoffs associated with individual entries. </p>\n<p>To create the split, we assigned each <code>pdb_id</code> a date based on the oldest <code>temporal_cutoff</code> within its MMseqs2 cluster (at 30% identity). For example, if two entries—PDB1 (2020-01-01) and PDB2 (2025-12-17) belong to the same cluster, both will be assigned to the training set because the cluster's cutoff becomes 2020-01-01. We chose this method to minimize the presence of structures in the validation set that are too similar to the training data and reduce the leakage due to presence of possible homologs. </p>\n<p>You can reproduce this procedure using the metadata in the extra folder:</p>\n<pre><code>import pandas as pd\n\n# Read the chain metadata \nmetadata = pd.read_csv('extra/rna_metadata.csv', parse_dates=True)\n\n# Assign the oldest temporal cutoff for each cluster to all members of the cluster\ncluster_temporal_cutoff = metadata.groupby('mmseqs_0.300')['temporal_cutoff'].min()\ncluster_temporal_cutoff.name = 'cluster_temporal_cutoff'\nmetadata_with_cluster_cutoff = metadata.merge(cluster_temporal_cutoff, on='mmseqs_0.300')\n# For each PDB ID, get the minimum cluster temporal cutoff among its chains\npdbid_split_cutoff = metadata_with_cluster_cutoff.groupby('pdb_id')['cluster_temporal_cutoff'].min()\n\n# Split the PDB IDs into training and validation sets based on the cluster temporal cutoff\n# NOTE: the sequences undergo further filtering, so we have supersets of PDB IDs here\ntrain_superset = set(pdbid_split_cutoff[pdbid_split_cutoff &lt; '2025-05-29' ].index)\nvalidation_superset = set(pdbid_split_cutoff[pdbid_split_cutoff &gt;= '2025-05-29'].index)\n\ntrain_sequences = pd.read_csv('train_sequences.csv')\nvalidation_sequences = pd.read_csv('validation_sequences.csv')\n\ntrain_set = set(train_sequences['target_id'])\nvalidation_set = set(validation_sequences['target_id'])\n\n# Validate\nassert validation_set.isdisjoint(train_superset)\nassert train_set.isdisjoint(validation_superset)\nassert train_set.issubset(train_superset)\nassert validation_set.issubset(validation_superset)\n</code></pre>\n<p>Hope that clarifies the issue.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3408605,
          "author_name": "WJui",
          "author_url": "",
          "post_date": "2026-02-21T00:27:51.767000",
          "content": "<p>Thank you for your clarification!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3408505": "Thank you for raising this issue. We prepared the train/validation split using clustering at a 30% identity level and a temporal cutoff for each cluster. Since the initial description of the data preparation step was unclear, we will update it shortly. The `temporal_cutoff`  in the `train_sequences.csv` and `validation_sequences.csv` are the cutoffs associated with individual entries. \n\nTo create the split, we assigned each `pdb_id` a date based on the oldest `temporal_cutoff` within its MMseqs2 cluster (at 30% identity). For example, if two entries—PDB1 (2020-01-01) and PDB2 (2025-12-17) belong to the same cluster, both will be assigned to the training set because the cluster's cutoff becomes 2020-01-01. We chose this method to minimize the presence of structures in the validation set that are too similar to the training data and reduce the leakage due to presence of possible homologs. \n\nYou can reproduce this procedure using the metadata in the extra folder:\n\n```python\nimport pandas as pd\n\n# Read the chain metadata \nmetadata = pd.read_csv('extra/rna_metadata.csv', parse_dates=True)\n\n# Assign the oldest temporal cutoff for each cluster to all members of the cluster\ncluster_temporal_cutoff = metadata.groupby('mmseqs_0.300')['temporal_cutoff'].min()\ncluster_temporal_cutoff.name = 'cluster_temporal_cutoff'\nmetadata_with_cluster_cutoff = metadata.merge(cluster_temporal_cutoff, on='mmseqs_0.300')\n# For each PDB ID, get the minimum cluster temporal cutoff among its chains\npdbid_split_cutoff = metadata_with_cluster_cutoff.groupby('pdb_id')['cluster_temporal_cutoff'].min()\n\n# Split the PDB IDs into training and validation sets based on the cluster temporal cutoff\n# NOTE: the sequences undergo further filtering, so we have supersets of PDB IDs here\ntrain_superset = set(pdbid_split_cutoff[pdbid_split_cutoff < '2025-05-29' ].index)\nvalidation_superset = set(pdbid_split_cutoff[pdbid_split_cutoff >= '2025-05-29'].index)\n\ntrain_sequences = pd.read_csv('train_sequences.csv')\nvalidation_sequences = pd.read_csv('validation_sequences.csv')\n\ntrain_set = set(train_sequences['target_id'])\nvalidation_set = set(validation_sequences['target_id'])\n\n# Validate\nassert validation_set.isdisjoint(train_superset)\nassert train_set.isdisjoint(validation_superset)\nassert train_set.issubset(train_superset)\nassert validation_set.issubset(validation_superset)\n```\nHope that clarifies the issue.\n",
    "3408192": "According to the dataset description,\n>The train_sequences.csv has not been filtered for sequence redundancy and contains PDB entries released before May 29, 2025 ...\n\nHowever, this image shows that the latest cutoff date is 2025-12-17\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24942389%2F063b9119dacf0196a80ee6f5c1594a25%2F20260220093031.jpg?generation=1771551056907558&alt=media)"
  }
}