{
  "id": 508174,
  "title": "Questions about Seq2Seq",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/508174",
  "author_name": "Fernando Melo",
  "post_date": "2024-05-28T14:28:40.924000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Guys, I am trying to learn more about Seq2Seq, but I am struggling with several points. Here are some questions:</p>\n<p>1 - I have some knowledge of Autoencoders. Normally, in Autoencoders, the input of the decoder is the output of the encoder (called the latent representation).</p>\n<p>But in Seq2Seq, I did some research on Google (and ChatGPT), and it seems that the decoder is usually fed with the previous state. Is this due to the nature of RNN/LSTM? Like in the code below:</p>\n<pre><code> :\n    def :\n        super(Seq2Seq, self).\n\n        self.encoder = nn.\n        self.decoder = nn.\n        self.fc = nn.\n\n    def forward(self, encoder_input, decoder_input):\n        # Codificador\n        _, (hidden, cell) = self.encoder(encoder_input)\n\n        # Decodificador\n        decoder_output, _ = self.decoder(decoder_input, (hidden, cell))\n\n        # Previsão\n        predictions = self.fc(decoder_output)\n        return predictions\n\n\n epoch  range(num_epochs):\n     encoder_input, target  train_loader:\n        # Inicializando o estado oculto\n        decoder_input = torch.zeros(batch_size, target.size(), output_dim)\n\n        # Forward pass\n        outputs = model(encoder_input, decoder_input)\n</code></pre>\n<p>It makes no sense to me to always start the decoder input with 0 (the code above was built by ChatGPT).</p>\n<p>2 - Strictly speaking about the shape of the data that will be fed to the model, should the correct shape be (batch_size, 60, 14) or (batch_size, 14, 60) (and (batch_size, 60, 23) for the outputs)? The first one looks correct to me, but I want to confirm.</p>\n<p>3 - Is Unet a type of Seq2Seq? I asked ChatGPT and it said that it is not Seq2Seq.</p>\n<p>4 - Can working with bottleneck architectures cause problems with the decoder part due to the \"reconstruction\" of the data? Can using techniques like output_padding to match the shapes cause problems?</p>\n<p>If someone can share a simple Seq2Seq model for learning purposes, it would be really helpful.</p>",
  "messages": [
    {
      "id": 2843265,
      "postDate": "2024-05-29T13:38:59.100Z",
      "content": "<p>I just made this baseline seq2seq notebook public: <a href=\"https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq\" target=\"_blank\">https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq</a></p>",
      "rawMarkdown": "I just made this baseline seq2seq notebook public: https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq",
      "votes": 5,
      "replies": [
        {
          "id": 2843370,
          "postDate": "2024-05-29T14:35:29.083Z",
          "content": "<p>Good job. A simple model with a lot of room for improvement, yet with a good score. I'm sure it will help a lot of people.</p>",
          "rawMarkdown": "Good job. A simple model with a lot of room for improvement, yet with a good score. I'm sure it will help a lot of people.",
          "votes": 2
        },
        {
          "id": 2845503,
          "postDate": "2024-05-30T15:17:27.547Z",
          "content": "<p>THX, really helpful. I am learning a lot from you.</p>\n<p>I'm studying your code, where does this architecture come from?<br>\nIsn't this a complete seq2seq in the sense, that you don't have an explicit encoder-decoder architecture?<br>\nAnd you don't have the input shape (batch, seq_len, features).</p>",
          "rawMarkdown": "THX, really helpful. I am learning a lot from you.\n\nI'm studying your code, where does this architecture come from?\nIsn't this a complete seq2seq in the sense, that you don't have an explicit encoder-decoder architecture?\nAnd you don't have the input shape (batch, seq_len, features).",
          "replies": [
            {
              "id": 2845738,
              "postDate": "2024-05-30T16:55:05.157Z",
              "content": "<p>You're confusing seq2seq (a general term) with causal attention (encoder-decoder).<br>\nThe notebook above implements seq2seq with a 1d cnn.</p>",
              "rawMarkdown": "You're confusing seq2seq (a general term) with causal attention (encoder-decoder).\nThe notebook above implements seq2seq with a 1d cnn.",
              "votes": 2
            },
            {
              "id": 2845772,
              "postDate": "2024-05-30T17:13:43.947Z",
              "content": "<p>The <code>x_to_seq</code> function reshape the \"flat\" input in a sequence of shape (Batch, 60, 25). In that case we don't have time/causal order in the sequence so we don't need to mask data from the future like is possible to do with classical encoder/decoder architecture, so we can use convolutions, (unmasked) transformers and/or bidirectional LSTM (or maybe some other fancy idea)</p>",
              "rawMarkdown": "The `x_to_seq` function reshape the \"flat\" input in a sequence of shape (Batch, 60, 25). In that case we don't have time/causal order in the sequence so we don't need to mask data from the future like is possible to do with classical encoder/decoder architecture, so we can use convolutions, (unmasked) transformers and/or bidirectional LSTM (or maybe some other fancy idea)"
            }
          ]
        },
        {
          "id": 2856785,
          "postDate": "2024-06-05T13:49:40.543Z",
          "content": "<p>Amedeo, I studied your notebook and our team is improving ours based on yours.</p>\n<p>I have one more question, how do you initially find this network architecture? Trial and error? Based on some literature network? Domain knowledge?</p>",
          "rawMarkdown": "Amedeo, I studied your notebook and our team is improving ours based on yours.\n\nI have one more question, how do you initially find this network architecture? Trial and error? Based on some literature network? Domain knowledge?"
        }
      ]
    },
    {
      "id": 2841765,
      "postDate": "2024-05-28T17:26:31.863Z",
      "content": "<p>'feed previous state' is when you use encoder-decoder. This is for problems where the output length is independent of the input length. For example, your input is a sentence at length X, and the output should be a translation of length Y.<br>\nBut here, the output should be the same length as the input, i.e., 60. So, there is no need for an encoder-decoder scheme, only an encoder. Hence, there is also no need to 'feed the previous state.' Unless you use an RNN/LSTM, of course, which works on the 'previous state' by design, but if your encoder is transformer or 1d unet etc. you just map 60 to 60. This is the 'seq to seq'. For example, consider the simplest model of 1 layer of 1d conv with 'same' paddings. Well, it might be challenging to build a good model of the sort if you don't have previous experience, there are some tricks to the trade haha. Since no one shared a good baseline, this competition is really not easy to get into.</p>",
      "rawMarkdown": "'feed previous state' is when you use encoder-decoder. This is for problems where the output length is independent of the input length. For example, your input is a sentence at length X, and the output should be a translation of length Y.\nBut here, the output should be the same length as the input, i.e., 60. So, there is no need for an encoder-decoder scheme, only an encoder. Hence, there is also no need to 'feed the previous state.' Unless you use an RNN/LSTM, of course, which works on the 'previous state' by design, but if your encoder is transformer or 1d unet etc. you just map 60 to 60. This is the 'seq to seq'. For example, consider the simplest model of 1 layer of 1d conv with 'same' paddings. Well, it might be challenging to build a good model of the sort if you don't have previous experience, there are some tricks to the trade haha. Since no one shared a good baseline, this competition is really not easy to get into.",
      "votes": 4
    },
    {
      "id": 2841332,
      "postDate": "2024-05-28T14:28:40.923Z",
      "content": "<p>Guys, I am trying to learn more about Seq2Seq, but I am struggling with several points. Here are some questions:</p>\n<p>1 - I have some knowledge of Autoencoders. Normally, in Autoencoders, the input of the decoder is the output of the encoder (called the latent representation).</p>\n<p>But in Seq2Seq, I did some research on Google (and ChatGPT), and it seems that the decoder is usually fed with the previous state. Is this due to the nature of RNN/LSTM? Like in the code below:</p>\n<pre><code> :\n    def :\n        super(Seq2Seq, self).\n\n        self.encoder = nn.\n        self.decoder = nn.\n        self.fc = nn.\n\n    def forward(self, encoder_input, decoder_input):\n        # Codificador\n        _, (hidden, cell) = self.encoder(encoder_input)\n\n        # Decodificador\n        decoder_output, _ = self.decoder(decoder_input, (hidden, cell))\n\n        # Previsão\n        predictions = self.fc(decoder_output)\n        return predictions\n\n\n epoch  range(num_epochs):\n     encoder_input, target  train_loader:\n        # Inicializando o estado oculto\n        decoder_input = torch.zeros(batch_size, target.size(), output_dim)\n\n        # Forward pass\n        outputs = model(encoder_input, decoder_input)\n</code></pre>\n<p>It makes no sense to me to always start the decoder input with 0 (the code above was built by ChatGPT).</p>\n<p>2 - Strictly speaking about the shape of the data that will be fed to the model, should the correct shape be (batch_size, 60, 14) or (batch_size, 14, 60) (and (batch_size, 60, 23) for the outputs)? The first one looks correct to me, but I want to confirm.</p>\n<p>3 - Is Unet a type of Seq2Seq? I asked ChatGPT and it said that it is not Seq2Seq.</p>\n<p>4 - Can working with bottleneck architectures cause problems with the decoder part due to the \"reconstruction\" of the data? Can using techniques like output_padding to match the shapes cause problems?</p>\n<p>If someone can share a simple Seq2Seq model for learning purposes, it would be really helpful.</p>",
      "rawMarkdown": "Guys, I am trying to learn more about Seq2Seq, but I am struggling with several points. Here are some questions:\n\n1 - I have some knowledge of Autoencoders. Normally, in Autoencoders, the input of the decoder is the output of the encoder (called the latent representation).\n\nBut in Seq2Seq, I did some research on Google (and ChatGPT), and it seems that the decoder is usually fed with the previous state. Is this due to the nature of RNN/LSTM? Like in the code below:\n\n```\nclass Seq2Seq(nn.Module):\n    def __init__(self, input_dim, output_dim, hidden_dim, num_layers):\n        super(Seq2Seq, self).__init__()\n        \n        self.encoder = nn.LSTM(input_dim, hidden_dim, num_layers, batch_first=True)\n        self.decoder = nn.LSTM(output_dim, hidden_dim, num_layers, batch_first=True)\n        self.fc = nn.Linear(hidden_dim, output_dim)\n        \n    def forward(self, encoder_input, decoder_input):\n        # Codificador\n        _, (hidden, cell) = self.encoder(encoder_input)\n        \n        # Decodificador\n        decoder_output, _ = self.decoder(decoder_input, (hidden, cell))\n        \n        # Previsão\n        predictions = self.fc(decoder_output)\n        return predictions\n\n\nfor epoch in range(num_epochs):\n    for encoder_input, target in train_loader:\n        # Inicializando o estado oculto\n        decoder_input = torch.zeros(batch_size, target.size(1), output_dim)\n        \n        # Forward pass\n        outputs = model(encoder_input, decoder_input)\n```\n\nIt makes no sense to me to always start the decoder input with 0 (the code above was built by ChatGPT).\n\n2 - Strictly speaking about the shape of the data that will be fed to the model, should the correct shape be (batch_size, 60, 14) or (batch_size, 14, 60) (and (batch_size, 60, 23) for the outputs)? The first one looks correct to me, but I want to confirm.\n\n3 - Is Unet a type of Seq2Seq? I asked ChatGPT and it said that it is not Seq2Seq.\n\n4 - Can working with bottleneck architectures cause problems with the decoder part due to the \"reconstruction\" of the data? Can using techniques like output_padding to match the shapes cause problems?\n\nIf someone can share a simple Seq2Seq model for learning purposes, it would be really helpful.\n\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2843265,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-05-29T13:38:59.100000",
      "content": "<p>I just made this baseline seq2seq notebook public: <a href=\"https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq\" target=\"_blank\">https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2843370,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-05-29T14:35:29.083000",
          "content": "<p>Good job. A simple model with a lot of room for improvement, yet with a good score. I'm sure it will help a lot of people.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2845503,
          "author_name": "Fernando Melo",
          "author_url": "",
          "post_date": "2024-05-30T15:17:27.547000",
          "content": "<p>THX, really helpful. I am learning a lot from you.</p>\n<p>I'm studying your code, where does this architecture come from?<br>\nIsn't this a complete seq2seq in the sense, that you don't have an explicit encoder-decoder architecture?<br>\nAnd you don't have the input shape (batch, seq_len, features).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2845738,
              "author_name": "sroger",
              "author_url": "",
              "post_date": "2024-05-30T16:55:05.157000",
              "content": "<p>You're confusing seq2seq (a general term) with causal attention (encoder-decoder).<br>\nThe notebook above implements seq2seq with a 1d cnn.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2845772,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-05-30T17:13:43.947000",
              "content": "<p>The <code>x_to_seq</code> function reshape the \"flat\" input in a sequence of shape (Batch, 60, 25). In that case we don't have time/causal order in the sequence so we don't need to mask data from the future like is possible to do with classical encoder/decoder architecture, so we can use convolutions, (unmasked) transformers and/or bidirectional LSTM (or maybe some other fancy idea)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2856785,
          "author_name": "Fernando Melo",
          "author_url": "",
          "post_date": "2024-06-05T13:49:40.543000",
          "content": "<p>Amedeo, I studied your notebook and our team is improving ours based on yours.</p>\n<p>I have one more question, how do you initially find this network architecture? Trial and error? Based on some literature network? Domain knowledge?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2841765,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-05-28T17:26:31.863000",
      "content": "<p>'feed previous state' is when you use encoder-decoder. This is for problems where the output length is independent of the input length. For example, your input is a sentence at length X, and the output should be a translation of length Y.<br>\nBut here, the output should be the same length as the input, i.e., 60. So, there is no need for an encoder-decoder scheme, only an encoder. Hence, there is also no need to 'feed the previous state.' Unless you use an RNN/LSTM, of course, which works on the 'previous state' by design, but if your encoder is transformer or 1d unet etc. you just map 60 to 60. This is the 'seq to seq'. For example, consider the simplest model of 1 layer of 1d conv with 'same' paddings. Well, it might be challenging to build a good model of the sort if you don't have previous experience, there are some tricks to the trade haha. Since no one shared a good baseline, this competition is really not easy to get into.</p>",
      "votes": 4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2843265": "I just made this baseline seq2seq notebook public: https://www.kaggle.com/code/abiolatti/keras-baseline-seq2seq",
    "2841765": "'feed previous state' is when you use encoder-decoder. This is for problems where the output length is independent of the input length. For example, your input is a sentence at length X, and the output should be a translation of length Y.\nBut here, the output should be the same length as the input, i.e., 60. So, there is no need for an encoder-decoder scheme, only an encoder. Hence, there is also no need to 'feed the previous state.' Unless you use an RNN/LSTM, of course, which works on the 'previous state' by design, but if your encoder is transformer or 1d unet etc. you just map 60 to 60. This is the 'seq to seq'. For example, consider the simplest model of 1 layer of 1d conv with 'same' paddings. Well, it might be challenging to build a good model of the sort if you don't have previous experience, there are some tricks to the trade haha. Since no one shared a good baseline, this competition is really not easy to get into.",
    "2841332": "Guys, I am trying to learn more about Seq2Seq, but I am struggling with several points. Here are some questions:\n\n1 - I have some knowledge of Autoencoders. Normally, in Autoencoders, the input of the decoder is the output of the encoder (called the latent representation).\n\nBut in Seq2Seq, I did some research on Google (and ChatGPT), and it seems that the decoder is usually fed with the previous state. Is this due to the nature of RNN/LSTM? Like in the code below:\n\n```\nclass Seq2Seq(nn.Module):\n    def __init__(self, input_dim, output_dim, hidden_dim, num_layers):\n        super(Seq2Seq, self).__init__()\n        \n        self.encoder = nn.LSTM(input_dim, hidden_dim, num_layers, batch_first=True)\n        self.decoder = nn.LSTM(output_dim, hidden_dim, num_layers, batch_first=True)\n        self.fc = nn.Linear(hidden_dim, output_dim)\n        \n    def forward(self, encoder_input, decoder_input):\n        # Codificador\n        _, (hidden, cell) = self.encoder(encoder_input)\n        \n        # Decodificador\n        decoder_output, _ = self.decoder(decoder_input, (hidden, cell))\n        \n        # Previsão\n        predictions = self.fc(decoder_output)\n        return predictions\n\n\nfor epoch in range(num_epochs):\n    for encoder_input, target in train_loader:\n        # Inicializando o estado oculto\n        decoder_input = torch.zeros(batch_size, target.size(1), output_dim)\n        \n        # Forward pass\n        outputs = model(encoder_input, decoder_input)\n```\n\nIt makes no sense to me to always start the decoder input with 0 (the code above was built by ChatGPT).\n\n2 - Strictly speaking about the shape of the data that will be fed to the model, should the correct shape be (batch_size, 60, 14) or (batch_size, 14, 60) (and (batch_size, 60, 23) for the outputs)? The first one looks correct to me, but I want to confirm.\n\n3 - Is Unet a type of Seq2Seq? I asked ChatGPT and it said that it is not Seq2Seq.\n\n4 - Can working with bottleneck architectures cause problems with the decoder part due to the \"reconstruction\" of the data? Can using techniques like output_padding to match the shapes cause problems?\n\nIf someone can share a simple Seq2Seq model for learning purposes, it would be really helpful.\n\n"
  }
}