{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":118765,"databundleVersionId":16320058}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# My approach for RNA folding\nHi! While I very much think I won't be a prize contender I did find my approach was a bit different then the others so I will post this here.\n\nMy main idea was for this competition there were mainly 2 ways to go about to win the prize\n\n1. Increase the quality of the 5 samples obtained from TBM/protenix in the first place. I did have an idea for this approach but I sadly didn't finish this in time.\n2. Generate a lot of diverse samples and then pick 5 that is the most likely to have a good TBM.\n\nI mainly focused all my effort into doing number 2. To do this approach, I genereted around 3500 samples of candidate rna structures for the validation set(retrieving from just training set) and using the evaluation code, I modified it so for each of these candidates I get a TM score.\nThen, my idea was to feature engineer each candidate structure in order to predict this TM score at pretty high accuracy, which did work quite a bit even with 5 fold validation using a random forest where I got like 0.98 R^2 between predicted TM score vs actual TM score.\n\nThen, at least on validation set, I noticed the optimal strategy was getting 2 of this highest scoring predicted TM scores + 3 after clustering with those features I engineered above.\n\nThe main issue with this approach was that \n1. For this testing set there wasn't enough time for me to generate a lot of candidate structures. Like I don't think I was able to generate 10 structures of protenix within the 8h timeframe consistently\n2. At least in terms of performance on the test set, it seemed like getting very high quality TBM is consistently better(ex filtering to 50% similarity)\n3. I used a lot of diverse models like RNAPro, openfold, boltz, protenix but all these take significant time which honestly just TBM+protenix seemed to always perform better\nSo overall, my issue was that I think I focused on the wrong strategy. I should have devoted my time on improving the strucutes cheaply in the 8h runtime rather than scoring but very much looking forward to seeing the top solutions. My guess is it'll be a bit like RNAPro but with stoichem+ligand knowledge and adapted to longer sequence.","metadata":{}}]}