{
  "id": 411341,
  "title": "What kind of NN?",
  "url": "/competitions/asl-fingerspelling/discussion/411341",
  "author_name": "MISAEL C RIBEIRO",
  "post_date": "2023-05-18T17:51:59.540000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi there, </p>\n<p>I was wondering, what kind of neural network is capable to do this video/text learning?</p>",
  "messages": [
    {
      "id": 2266906,
      "postDate": "2023-05-20T13:05:29.680Z",
      "content": "<p>You can try a sequence to sequence transformer. Maybe use an existing translation architecture :) </p>",
      "rawMarkdown": "You can try a sequence to sequence transformer. Maybe use an existing translation architecture :) ",
      "votes": 1
    },
    {
      "id": 2265474,
      "postDate": "2023-05-19T09:09:14.610Z",
      "content": "<p>I could recommend a transformer encoder/decoder architecture.<br>\nIn contrast to conventional text2text models, the encoder will need to be adapted to process a video of coordinates.</p>\n<p>A <a href=\"https://keras.io/examples/nlp/neural_machine_translation_with_transformer/\" target=\"_blank\">detailed tutorial</a> is available in the Keras documentation showing how to build a transformer encoder/decoder from scratch to translate English to Spaninish.</p>",
      "rawMarkdown": "I could recommend a transformer encoder/decoder architecture.\nIn contrast to conventional text2text models, the encoder will need to be adapted to process a video of coordinates.\n\nA [detailed tutorial](https://keras.io/examples/nlp/neural_machine_translation_with_transformer/) is available in the Keras documentation showing how to build a transformer encoder/decoder from scratch to translate English to Spaninish.",
      "votes": 2
    },
    {
      "id": 2265311,
      "postDate": "2023-05-19T06:47:14.943Z",
      "content": "<p>On top of <a href=\"https://www.kaggle.com/ericka42\" target=\"_blank\">@ericka42</a> answer, I would like to add some notes regarding string prediction. Predicting characters (especially a string of unknown length) is not a straightforward task. Recently, I wrote some thoughts on combining CNN + RNN + CTC loss. If you are interested in more details, you could check out my post here - <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512\" target=\"_blank\">Sequence modeling with CTC loss</a></p>",
      "rawMarkdown": "On top of @ericka42 answer, I would like to add some notes regarding string prediction. Predicting characters (especially a string of unknown length) is not a straightforward task. Recently, I wrote some thoughts on combining CNN + RNN + CTC loss. If you are interested in more details, you could check out my post here - [Sequence modeling with CTC loss](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512)",
      "votes": 2,
      "replies": [
        {
          "id": 2265646,
          "postDate": "2023-05-19T12:23:05.220Z",
          "content": "<p>When it comes to NNs, i have a great time understanding how everything works, but the coding part is really hard. Thanks for the tip! CTC Loss seems to do a good jobo indeed</p>",
          "rawMarkdown": "When it comes to NNs, i have a great time understanding how everything works, but the coding part is really hard. Thanks for the tip! CTC Loss seems to do a good jobo indeed",
          "votes": 1,
          "replies": [
            {
              "id": 2265676,
              "postDate": "2023-05-19T12:52:49.267Z",
              "content": "<p>I understand you. To be honest, the same is true for me. I have limited practice, so sometimes I have a brilliant idea and I am not able to implement it in a reasonable time. We need to practice in order to improve our skills. </p>\n<p>In case you would fine it helpful, I will share a <a href=\"https://github.com/Abto-Coursework-Mentoring-2020/ocr-engine/blob/main/training/single_line_model_training.py\" target=\"_blank\">link</a> to the notebook with OCR mode (image to text) developed by one of my students. It is not perfect, but it may be a baseline. </p>",
              "rawMarkdown": "I understand you. To be honest, the same is true for me. I have limited practice, so sometimes I have a brilliant idea and I am not able to implement it in a reasonable time. We need to practice in order to improve our skills. \n\nIn case you would fine it helpful, I will share a [link](https://github.com/Abto-Coursework-Mentoring-2020/ocr-engine/blob/main/training/single_line_model_training.py) to the notebook with OCR mode (image to text) developed by one of my students. It is not perfect, but it may be a baseline. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2264989,
      "postDate": "2023-05-18T22:55:11.073Z",
      "content": "<p>For video/text learning tasks, a type of neural network that is commonly used is a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This combination allows the network to effectively process both the spatial information in videos and the sequential nature of text data.</p>\n<p>Here's an overview of how CNNs and RNNs are typically used in video/text learning:</p>\n<p>Convolutional Neural Networks (CNNs): CNNs are well-suited for processing visual information, making them ideal for analyzing video frames. They excel at capturing spatial patterns and features in images. In video learning, CNNs are often used to extract visual features from individual frames or small temporal windows (e.g., consecutive frames).</p>\n<p>Recurrent Neural Networks (RNNs): RNNs are designed to handle sequential data, making them suitable for processing text and capturing dependencies over time. RNNs have a \"memory\" that allows them to retain information from previous steps and use it to inform predictions at each time step. In text learning, RNNs can process sequences of words or characters, capturing contextual information and modeling the sequential nature of language.</p>\n<p>Combination of CNNs and RNNs: To tackle video/text learning tasks, the two networks can be combined in different ways. One common approach is to use a CNN as a feature extractor for video frames, and then pass the extracted features as input to an RNN for temporal modeling. This allows the network to capture both the spatial information in individual frames and the temporal dependencies between frames.</p>\n<p>In recent years, there have been advancements in architectures specifically designed for video/text learning, such as 3D CNNs, which capture both spatial and temporal features simultaneously, and Transformer models, which have shown great success in natural language processing tasks. These architectures often incorporate elements from both CNNs and RNNs to effectively handle the complexity of video and text data.</p>\n<p>It's important to note that the specific choice of neural network architecture depends on the specific video/text learning task at hand, the available data, and the desired performance. Researchers and practitioners often experiment with different architectures and variations to find the best approach for a particular problem.</p>",
      "rawMarkdown": "For video/text learning tasks, a type of neural network that is commonly used is a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This combination allows the network to effectively process both the spatial information in videos and the sequential nature of text data.\n\nHere's an overview of how CNNs and RNNs are typically used in video/text learning:\n\nConvolutional Neural Networks (CNNs): CNNs are well-suited for processing visual information, making them ideal for analyzing video frames. They excel at capturing spatial patterns and features in images. In video learning, CNNs are often used to extract visual features from individual frames or small temporal windows (e.g., consecutive frames).\n\nRecurrent Neural Networks (RNNs): RNNs are designed to handle sequential data, making them suitable for processing text and capturing dependencies over time. RNNs have a \"memory\" that allows them to retain information from previous steps and use it to inform predictions at each time step. In text learning, RNNs can process sequences of words or characters, capturing contextual information and modeling the sequential nature of language.\n\nCombination of CNNs and RNNs: To tackle video/text learning tasks, the two networks can be combined in different ways. One common approach is to use a CNN as a feature extractor for video frames, and then pass the extracted features as input to an RNN for temporal modeling. This allows the network to capture both the spatial information in individual frames and the temporal dependencies between frames.\n\nIn recent years, there have been advancements in architectures specifically designed for video/text learning, such as 3D CNNs, which capture both spatial and temporal features simultaneously, and Transformer models, which have shown great success in natural language processing tasks. These architectures often incorporate elements from both CNNs and RNNs to effectively handle the complexity of video and text data.\n\nIt's important to note that the specific choice of neural network architecture depends on the specific video/text learning task at hand, the available data, and the desired performance. Researchers and practitioners often experiment with different architectures and variations to find the best approach for a particular problem.",
      "votes": 2,
      "replies": [
        {
          "id": 2265071,
          "postDate": "2023-05-19T00:51:36.357Z",
          "content": "<p>What an amazing answer! Thanks! Is there any notebook that can I learn about combining two different NN like you said? Is there some notebook similar about video/text learning tasks in this competition?</p>",
          "rawMarkdown": "What an amazing answer! Thanks! Is there any notebook that can I learn about combining two different NN like you said? Is there some notebook similar about video/text learning tasks in this competition?"
        }
      ]
    },
    {
      "id": 2264783,
      "postDate": "2023-05-18T17:51:59.540Z",
      "content": "<p>Hi there, </p>\n<p>I was wondering, what kind of neural network is capable to do this video/text learning?</p>",
      "rawMarkdown": "Hi there, \n\nI was wondering, what kind of neural network is capable to do this video/text learning?"
    }
  ],
  "comments": [
    {
      "id": 2266906,
      "author_name": "white lady",
      "author_url": "",
      "post_date": "2023-05-20T13:05:29.680000",
      "content": "<p>You can try a sequence to sequence transformer. Maybe use an existing translation architecture :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2265474,
      "author_name": "Mark Wijkhuizen",
      "author_url": "",
      "post_date": "2023-05-19T09:09:14.610000",
      "content": "<p>I could recommend a transformer encoder/decoder architecture.<br>\nIn contrast to conventional text2text models, the encoder will need to be adapted to process a video of coordinates.</p>\n<p>A <a href=\"https://keras.io/examples/nlp/neural_machine_translation_with_transformer/\" target=\"_blank\">detailed tutorial</a> is available in the Keras documentation showing how to build a transformer encoder/decoder from scratch to translate English to Spaninish.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2265311,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-05-19T06:47:14.943000",
      "content": "<p>On top of <a href=\"https://www.kaggle.com/ericka42\" target=\"_blank\">@ericka42</a> answer, I would like to add some notes regarding string prediction. Predicting characters (especially a string of unknown length) is not a straightforward task. Recently, I wrote some thoughts on combining CNN + RNN + CTC loss. If you are interested in more details, you could check out my post here - <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512\" target=\"_blank\">Sequence modeling with CTC loss</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2265646,
          "author_name": "MISAEL C RIBEIRO",
          "author_url": "",
          "post_date": "2023-05-19T12:23:05.220000",
          "content": "<p>When it comes to NNs, i have a great time understanding how everything works, but the coding part is really hard. Thanks for the tip! CTC Loss seems to do a good jobo indeed</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2265676,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-05-19T12:52:49.267000",
              "content": "<p>I understand you. To be honest, the same is true for me. I have limited practice, so sometimes I have a brilliant idea and I am not able to implement it in a reasonable time. We need to practice in order to improve our skills. </p>\n<p>In case you would fine it helpful, I will share a <a href=\"https://github.com/Abto-Coursework-Mentoring-2020/ocr-engine/blob/main/training/single_line_model_training.py\" target=\"_blank\">link</a> to the notebook with OCR mode (image to text) developed by one of my students. It is not perfect, but it may be a baseline. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2264989,
      "author_name": "Ericka42",
      "author_url": "",
      "post_date": "2023-05-18T22:55:11.073000",
      "content": "<p>For video/text learning tasks, a type of neural network that is commonly used is a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This combination allows the network to effectively process both the spatial information in videos and the sequential nature of text data.</p>\n<p>Here's an overview of how CNNs and RNNs are typically used in video/text learning:</p>\n<p>Convolutional Neural Networks (CNNs): CNNs are well-suited for processing visual information, making them ideal for analyzing video frames. They excel at capturing spatial patterns and features in images. In video learning, CNNs are often used to extract visual features from individual frames or small temporal windows (e.g., consecutive frames).</p>\n<p>Recurrent Neural Networks (RNNs): RNNs are designed to handle sequential data, making them suitable for processing text and capturing dependencies over time. RNNs have a \"memory\" that allows them to retain information from previous steps and use it to inform predictions at each time step. In text learning, RNNs can process sequences of words or characters, capturing contextual information and modeling the sequential nature of language.</p>\n<p>Combination of CNNs and RNNs: To tackle video/text learning tasks, the two networks can be combined in different ways. One common approach is to use a CNN as a feature extractor for video frames, and then pass the extracted features as input to an RNN for temporal modeling. This allows the network to capture both the spatial information in individual frames and the temporal dependencies between frames.</p>\n<p>In recent years, there have been advancements in architectures specifically designed for video/text learning, such as 3D CNNs, which capture both spatial and temporal features simultaneously, and Transformer models, which have shown great success in natural language processing tasks. These architectures often incorporate elements from both CNNs and RNNs to effectively handle the complexity of video and text data.</p>\n<p>It's important to note that the specific choice of neural network architecture depends on the specific video/text learning task at hand, the available data, and the desired performance. Researchers and practitioners often experiment with different architectures and variations to find the best approach for a particular problem.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2265071,
          "author_name": "MISAEL C RIBEIRO",
          "author_url": "",
          "post_date": "2023-05-19T00:51:36.357000",
          "content": "<p>What an amazing answer! Thanks! Is there any notebook that can I learn about combining two different NN like you said? Is there some notebook similar about video/text learning tasks in this competition?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2266906": "You can try a sequence to sequence transformer. Maybe use an existing translation architecture :) ",
    "2265474": "I could recommend a transformer encoder/decoder architecture.\nIn contrast to conventional text2text models, the encoder will need to be adapted to process a video of coordinates.\n\nA [detailed tutorial](https://keras.io/examples/nlp/neural_machine_translation_with_transformer/) is available in the Keras documentation showing how to build a transformer encoder/decoder from scratch to translate English to Spaninish.",
    "2265311": "On top of @ericka42 answer, I would like to add some notes regarding string prediction. Predicting characters (especially a string of unknown length) is not a straightforward task. Recently, I wrote some thoughts on combining CNN + RNN + CTC loss. If you are interested in more details, you could check out my post here - [Sequence modeling with CTC loss](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/409512)",
    "2264989": "For video/text learning tasks, a type of neural network that is commonly used is a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This combination allows the network to effectively process both the spatial information in videos and the sequential nature of text data.\n\nHere's an overview of how CNNs and RNNs are typically used in video/text learning:\n\nConvolutional Neural Networks (CNNs): CNNs are well-suited for processing visual information, making them ideal for analyzing video frames. They excel at capturing spatial patterns and features in images. In video learning, CNNs are often used to extract visual features from individual frames or small temporal windows (e.g., consecutive frames).\n\nRecurrent Neural Networks (RNNs): RNNs are designed to handle sequential data, making them suitable for processing text and capturing dependencies over time. RNNs have a \"memory\" that allows them to retain information from previous steps and use it to inform predictions at each time step. In text learning, RNNs can process sequences of words or characters, capturing contextual information and modeling the sequential nature of language.\n\nCombination of CNNs and RNNs: To tackle video/text learning tasks, the two networks can be combined in different ways. One common approach is to use a CNN as a feature extractor for video frames, and then pass the extracted features as input to an RNN for temporal modeling. This allows the network to capture both the spatial information in individual frames and the temporal dependencies between frames.\n\nIn recent years, there have been advancements in architectures specifically designed for video/text learning, such as 3D CNNs, which capture both spatial and temporal features simultaneously, and Transformer models, which have shown great success in natural language processing tasks. These architectures often incorporate elements from both CNNs and RNNs to effectively handle the complexity of video and text data.\n\nIt's important to note that the specific choice of neural network architecture depends on the specific video/text learning task at hand, the available data, and the desired performance. Researchers and practitioners often experiment with different architectures and variations to find the best approach for a particular problem.",
    "2264783": "Hi there, \n\nI was wondering, what kind of neural network is capable to do this video/text learning?"
  }
}