2026/07/02 更新

写真a

ツキヤマ ショウ
築山 翔
TSUKIYAMA SHO
所属
生命理工学院 助教
職名
助教

論文

  • RNArefine: AI-guided Atomic-Level Refinement of RNA Structures

    Sho Tsukiyama, Yang Li, Kengo Sato, Hiroyuki Kurata, Yang Zhang

    2026年6月

  • DRfold2 is a deep learning-based tool that enables efficient and accurate RNA structure prediction

    Yang Li, Chenjie Feng, Xi Zhang, Sho Tsukiyama, Duanyu Feng, Yang Zhang

    PLOS Biology   24 ( 2 )   e3003659 - e3003659   2026年2月

     詳細を見る

    掲載種別:研究論文(学術雑誌)   出版者・発行元:Public Library of Science (PLoS)  

    RNA structures are essential for understanding their biological functions and developing RNA-targeted therapeutics. However, accurate RNA structure prediction from sequence remains a crucial challenge. We introduce DRfold2, a deep learning framework that integrates a novel pre-trained RNA Composite Language Model (RCLM) with a denoising structure module for end-to-end RNA structure prediction. Based solely on single sequence, DRfold2 achieves superior performance in both global topology and secondary structure predictions over other state-of-the-art approaches across multiple benchmark tests from diverse species. Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, with a more than 100% increase in contact prediction precision compared to existing methods. Furthermore, DRfold2 demonstrates high complementarity with AlphaFold3, achieving statistically significant accuracy gains when integrated into our optimization framework. By uniquely combining composite language modeling, denoising-based end-to-end learning, and deep learning-guided post-optimization, DRfold2 establishes a distinct direction for advancing ab initio RNA structure prediction.

    DOI: 10.1371/journal.pbio.3003659

    researchmap

  • MLm5C: A high-precision human RNA 5-methylcytosine sites predictor based on a combination of hybrid machine learning models. 国際誌

    Hiroyuki Kurata, Md Harun-Or-Roshid, Md Mehedi Hasan, Sho Tsukiyama, Kazuhiro Maeda, Balachandran Manavalan

    Methods (San Diego, Calif.)   227   37 - 47   2024年7月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    RNA modification serves as a pivotal component in numerous biological processes. Among the prevalent modifications, 5-methylcytosine (m5C) significantly influences mRNA export, translation efficiency and cell differentiation and are also associated with human diseases, including Alzheimer's disease, autoimmune disease, cancer, and cardiovascular diseases. Identification of m5C is critically responsible for understanding the RNA modification mechanisms and the epigenetic regulation of associated diseases. However, the large-scale experimental identification of m5C present significant challenges due to labor intensity and time requirements. Several computational tools, using machine learning, have been developed to supplement experimental methods, but identifying these sites lack accuracy and efficiency. In this study, we introduce a new predictor, MLm5C, for precise prediction of m5C sites using sequence data. Briefly, we evaluated eleven RNA sequence-derived features with four basic machine learning algorithms to generate baseline models. From these 44 models, we ranked them based on their performance and subsequently stacked the Top 20 baseline models as the best model, named MLm5C. The MLm5C outperformed the-state-of-the-art predictors. Notably, the optimization of the sequence length surrounding the modification sites significantly improved the prediction performance. MLm5C is an invaluable tool in accelerating the detection of m5C sites within the human genome, thereby facilitating in the characterization of their roles in post-transcriptional regulation.

    DOI: 10.1016/j.ymeth.2024.05.004

    PubMed

    researchmap

  • PredIL13: Stacking a variety of machine and deep learning methods with ESM-2 language model for identifying IL13-inducing peptides. 国際誌

    Hiroyuki Kurata, Md Harun-Or-Roshid, Sho Tsukiyama, Kazuhiro Maeda

    PloS one   19 ( 8 )   e0309078   2024年

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    Interleukin (IL)-13 has emerged as one of the recently identified cytokine. Since IL-13 causes the severity of COVID-19 and alters crucial biological processes, it is urgent to explore novel molecules or peptides capable of including IL-13. Computational prediction has received attention as a complementary method to in-vivo and in-vitro experimental identification of IL-13 inducing peptides, because experimental identification is time-consuming, laborious, and expensive. A few computational tools have been presented, including the IL13Pred and iIL13Pred. To increase prediction capability, we have developed PredIL13, a cutting-edge ensemble learning method with the latest ESM-2 protein language model. This method stacked the probability scores outputted by 168 single-feature machine/deep learning models, and then trained a logistic regression-based meta-classifier with the stacked probability score vectors. The key technology was to implement ESM-2 and to select the optimal single-feature models according to their absolute weight coefficient for logistic regression (AWCLR), an indicator of the importance of each single-feature model. Especially, the sequential deletion of single-feature models based on the iterative AWCLR ranking (SDIWC) method constructed the meta-classifier consisting of the top 16 single-feature models, named PredIL13, while considering the model's accuracy. The PredIL13 greatly outperformed the-state-of-the-art predictors, thus is an invaluable tool for accelerating the detection of IL13-inducing peptide within the human genome.

    DOI: 10.1371/journal.pone.0309078

    PubMed

    researchmap

  • CNN6mA: Interpretable neural network model based on position-specific CNN and cross-interactive network for 6mA site prediction. 国際誌

    Sho Tsukiyama, Md Mehedi Hasan, Hiroyuki Kurata

    Computational and structural biotechnology journal   21   644 - 654   2023年

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    N6-methyladenine (6mA) plays a critical role in various epigenetic processing including DNA replication, DNA repair, silencing, transcription, and diseases such as cancer. To understand such epigenetic mechanisms, 6 mA has been detected by high-throughput technologies on a genome-wide scale at single-base resolution, together with conventional methods such as immunoprecipitation, mass spectrometry and capillary electrophoresis, but these experimental approaches are time-consuming and laborious. To complement these problems, we have developed a CNN-based 6 mA site predictor, named CNN6mA, which proposed two new architectures: a position-specific 1-D convolutional layer and a cross-interactive network. In the position-specific 1-D convolutional layer, position-specific filters with different window sizes were applied to an inquiry sequence instead of sharing the same filters over all positions in order to extract the position-specific features at different levels. The cross-interactive network explored the relationships between all the nucleotide patterns within the inquiry sequence. Consequently, CNN6mA outperformed the existing state-of-the-art models in many species and created the contribution score vector that intelligibly interpret the prediction mechanism. The source codes and web application in CNN6mA are freely accessible at https://github.com/kuratahiroyuki/CNN6mA.git and http://kurata35.bio.kyutech.ac.jp/CNN6mA/, respectively.

    DOI: 10.1016/j.csbj.2022.12.043

    PubMed

    researchmap

  • ICAN: interpretable cross-attention network for identifying drug and target protein interactions 国際誌

    Hiroyuki Kurata, Sho Tsukiyama

    PloS one   17 ( 10 )   e0276609   2022年8月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    Drug-target protein interaction (DTI) identification is fundamental for drug discovery and drug repositioning, because therapeutic drugs act on disease-causing proteins. However, the DTI identification process often requires expensive and time-consuming tasks, including biological experiments involving large numbers of candidate compounds. Thus, a variety of computation approaches have been developed. Of the many approaches available, chemo-genomics feature-based methods have attracted considerable attention. These methods compute the feature descriptors of drugs and proteins as the input data to train machine and deep learning models to enable accurate prediction of unknown DTIs. In addition, attention-based learning methods have been proposed to identify and interpret DTI mechanisms. However, improvements are needed for enhancing prediction performance and DTI mechanism elucidation. To address these problems, we developed an attention-based method designated the interpretable cross-attention network (ICAN), which predicts DTIs using the Simplified Molecular Input Line Entry System of drugs and amino acid sequences of target proteins. We optimized the attention mechanism architecture by exploring the cross-attention or self-attention, attention layer depth, and selection of the context matrixes from the attention mechanism. We found that a plain attention mechanism that decodes drug-related protein context features without any protein-related drug context features effectively achieved high performance. The ICAN outperformed state-of-the-art methods in several metrics on the DAVIS dataset and first revealed with statistical significance that some weighted sites in the cross-attention weight matrix represent experimental binding sites, thus demonstrating the high interpretability of the results. The program is freely available at https://github.com/kuratahiroyuki/ICAN.

    DOI: 10.1101/2022.08.04.502877

    PubMed

    researchmap

  • Deepm5C: A deep-learning-based hybrid framework for identifying human RNA N5-methylcytosine sites using a stacking strategy. 国際誌

    Md Mehedi Hasan, Sho Tsukiyama, Jae Youl Cho, Hiroyuki Kurata, Md Ashad Alam, Xiaowen Liu, Balachandran Manavalan, Hong-Wen Deng

    Molecular therapy : the journal of the American Society of Gene Therapy   30 ( 8 )   2856 - 2867   2022年8月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    As one of the most prevalent post-transcriptional epigenetic modifications, N5-methylcytosine (m5C) plays an essential role in various cellular processes and disease pathogenesis. Therefore, it is important accurately identify m5C modifications in order to gain a deeper understanding of cellular processes and other possible functional mechanisms. Although a few computational methods have been proposed, their respective models have been developed using small training datasets. Hence, their practical application is quite limited in genome-wide detection. To overcome the existing limitations, we propose Deepm5C, a bioinformatics method for identifying RNA m5C sites throughout the human genome. To develop Deepm5C, we constructed a novel benchmarking dataset and investigated a mixture of three conventional feature-encoding algorithms and a feature derived from word-embedding approaches. Afterward, four variants of deep-learning classifiers and four commonly used conventional classifiers were employed and trained with the four encodings, ultimately obtaining 32 baseline models. A stacking strategy is effectively utilized by integrating the predicted output of the optimal baseline models and trained with a one-dimensional (1D) convolutional neural network. As a result, the Deepm5C predictor achieved excellent performance during cross-validation with a Matthews correlation coefficient and an accuracy of 0.697 and 0.855, respectively. The corresponding metrics during the independent test were 0.691 and 0.852, respectively. Overall, Deepm5C achieved a more accurate and stable performance than the baseline models and significantly outperformed the existing predictors, demonstrating the effectiveness of our proposed hybrid framework. Furthermore, Deepm5C is expected to assist community-wide efforts in identifying putative m5Cs and to formulate the novel testable biological hypothesis.

    DOI: 10.1016/j.ymthe.2022.05.001

    PubMed

    researchmap

  • iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model 国際誌

    Hiroyuki Kurata, Sho Tsukiyama, Balachandran Manavalan

    Briefings in Bioinformatics   23 ( 4 )   2022年7月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    The COVID-19 pandemic caused several million deaths worldwide. Development of anti-coronavirus drugs is thus urgent. Unlike conventional non-peptide drugs, antiviral peptide drugs are highly specific, easy to synthesize and modify, and not highly susceptible to drug resistance. To reduce the time and expense involved in screening thousands of peptides and assaying their antiviral activity, computational predictors for identifying anti-coronavirus peptides (ACVPs) are needed. However, few experimentally verified ACVP samples are available, even though a relatively large number of antiviral peptides (AVPs) have been discovered. In this study, we attempted to predict ACVPs using an AVP dataset and a small collection of ACVPs. Using conventional features, a binary profile and a word-embedding word2vec (W2V), we systematically explored five different machine learning methods: Transformer, Convolutional Neural Network, bidirectional Long Short-Term Memory, Random Forest (RF) and Support Vector Machine. Via exhaustive searches, we found that the RF classifier with W2V consistently achieved better performance on different datasets. The two main controlling factors were: (i) the dataset-specific W2V dictionary was generated from the training and independent test datasets instead of the widely used general UniProt proteome and (ii) a systematic search was conducted and determined the optimal k-mer value in W2V, which provides greater discrimination between positive and negative samples. Therefore, our proposed method, named iACVP, consistently provides better prediction performance compared with existing state-of-the-art methods. To assist experimentalists in identifying putative ACVPs, we implemented our model as a web server accessible via the following link: http://kurata35.bio.kyutech.ac.jp/iACVP.

    DOI: 10.1093/bib/bbac265

    PubMed

    researchmap

  • Cross-attention PHV: Prediction of human and virus protein-protein interactions using cross-attention–based neural networks 国際誌

    Sho Tsukiyama, Hiroyuki Kurata

    Computational and structural biotechnology journal   20   5564 - 5573   2022年7月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    Viral infections represent a major health concern worldwide. The alarming rate at which SARS-CoV-2 spreads, for example, led to a worldwide pandemic. Viruses incorporate genetic material into the host genome to hijack host cell functions such as the cell cycle and apoptosis. In these viral processes, protein-protein interactions (PPIs) play critical roles. Therefore, the identification of PPIs between humans and viruses is crucial for understanding the infection mechanism and host immune responses to viral infections and for discovering effective drugs. Experimental methods including mass spectrometry-based proteomics and yeast two-hybrid assays are widely used to identify human-virus PPIs, but these experimental methods are time-consuming, expensive, and laborious. To overcome this problem, we developed a novel computational predictor, named cross-attention PHV, by implementing two key technologies of the cross-attention mechanism and a one-dimensional convolutional neural network (1D-CNN). The cross-attention mechanisms were very effective in enhancing prediction and generalization abilities. Application of 1D-CNN to the word2vec-generated feature matrices reduced computational costs, thus extending the allowable length of protein sequences to 9000 amino acid residues. Cross-attention PHV outperformed existing state-of-the-art models using a benchmark dataset and accurately predicted PPIs for unknown viruses. Cross-attention PHV also predicted human-SARS-CoV-2 PPIs with area under the curve values >0.95. The Cross-attention PHV web server and source codes are freely available at https://kurata35.bio.kyutech.ac.jp/Cross-attention_PHV/ and https://github.com/kuratahiroyuki/Cross-Attention_PHV, respectively.

    DOI: 10.1101/2022.07.03.498630

    PubMed

    researchmap

  • BERT6mA: prediction of DNA N6-methyladenine site using deep learning-based approaches 国際誌

    Sho Tsukiyama, Md Mehedi Hasan, Hong-Wen Deng, Hiroyuki Kurata

    Briefings in Bioinformatics   23 ( 2 )   2022年3月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)   出版者・発行元:Oxford University Press ({OUP})  

    N6-methyladenine (6mA) is associated with important roles in DNA replication, DNA repair, transcription, regulation of gene expression. Several experimental methods were used to identify DNA modifications. However, these experimental methods are costly and time-consuming. To detect the 6mA and complement these shortcomings of experimental methods, we proposed a novel, deep leaning approach called BERT6mA. To compare the BERT6mA with other deep learning approaches, we used the benchmark datasets including 11 species. The BERT6mA presented the highest AUCs in eight species in independent tests. Furthermore, BERT6mA showed higher and comparable performance with the state-of-the-art models while the BERT6mA showed poor performances in a few species with a small sample size. To overcome this issue, pretraining and fine-tuning between two species were applied to the BERT6mA. The pretrained and fine-tuned models on specific species presented higher performances than other models even for the species with a small sample size. In addition to the prediction, we analyzed the attention weights generated by BERT6mA to reveal how the BERT6mA model extracts critical features responsible for the 6mA prediction. To facilitate biological sciences, the BERT6mA online web server and its source codes are freely accessible at https://github.com/kuratahiroyuki/BERT6mA.git, respectively.

    DOI: 10.1093/bib/bbac053

    PubMed

    researchmap

  • LSTM-PHV: Prediction of human-virus protein-protein interactions by LSTM with word2vec 国際誌

    Sho Tsukiyama, Md Mehedi Hasan, Satoshi Fujii, Hiroyuki Kurata

    Briefings in bioinformatics   22 ( 6 )   2021年2月

     詳細を見る

    記述言語:英語   掲載種別:研究論文(学術雑誌)  

    Viral infection involves a large number of protein-protein interactions (PPIs) between human and virus. The PPIs range from the initial binding of viral coat proteins to host membrane receptors to the hijacking of host transcription machinery. However, few interspecies PPIs have been identified, because experimental methods including mass spectrometry are time-consuming and expensive, and molecular dynamic simulation is limited only to the proteins whose 3D structures are solved. Sequence-based machine learning methods are expected to overcome these problems. We have first developed the LSTM model with word2vec to predict PPIs between human and virus, named LSTM-PHV, by using amino acid sequences alone. The LSTM-PHV effectively learnt the training data with a highly imbalanced ratio of positive to negative samples and achieved AUCs of 0.976 and 0.973 and accuracies of 0.984 and 0.985 on the training and independent datasets, respectively. In predicting PPIs between human and unknown or new virus, the LSTM-PHV learned greatly outperformed the existing state-of-the-art PPI predictors. Interestingly, learning of only sequence contexts as words is sufficient for PPI prediction. Use of uniform manifold approximation and projection demonstrated that the LSTM-PHV clearly distinguished the positive PPI samples from the negative ones. We presented the LSTM-PHV online web server and support data that are freely available at http://kurata35.bio.kyutech.ac.jp/LSTM-PHV.

    DOI: 10.1101/2021.02.26.432975

    PubMed

    researchmap

▼全件表示

講演・口頭発表等

  • 深層学習とシミュレーションを用いたRNA3次元構造の予測

    築山 翔, 倉田 博之, 佐藤 健吾, Yang Zhang

    2025年日本バイオインフォマティクス学会年会  2025年9月 

     詳細を見る

    開催年月日: 2025年9月

    記述言語:英語   会議種別:ポスター発表  

    researchmap

  • フラグメントの重ね合わせと最適化に基づく粗視化RNA構造からの全原子構造の生成

    築山 翔, 倉田博之, 佐藤健吾, Yang Zhang

    情報処理学会 (バイオ情報学研究会)  2025年6月 

     詳細を見る

    開催年月日: 2025年6月

    記述言語:日本語   会議種別:口頭発表(一般)  

    researchmap

  • 深層学習によるRNA構造予測とRefinement手法の開発

    築山 翔, 倉田 博之, Yang Zhang

    情報処理学会 (バイオ情報学研究会)  2024年12月 

     詳細を見る

    開催年月日: 2024年12月

    記述言語:日本語   会議種別:口頭発表(一般)  

    researchmap

  • Prediction of RNA N5-methylcytosine sites by interpretable deep learning model

    Sho Tsukiyama, Hiroyuki Kurata

    2022年日本バイオインフォマティクス学会年会  2022年9月 

     詳細を見る

    開催年月日: 2022年9月

    記述言語:英語   会議種別:ポスター発表  

    researchmap

  • Attention機構ベースの深層学習によるヒト-ウイルスタンパク質間相互作用予測手法の提案

    築山 翔, 倉田 博之

    情報処理学会 (バイオ情報学研究会)  2022年3月 

     詳細を見る

    開催年月日: 2022年3月

    記述言語:日本語   会議種別:口頭発表(一般)  

    researchmap

  • BERTとword2vecを用いたN6-Methyladenineサイトの予測

    Sho Tsukiyama, Mehedi Md. Hasan, Hiroyuki Kurata

    2021年日本バイオインフォマティクス学会年会  2021年9月 

     詳細を見る

    開催年月日: 2021年9月

    記述言語:日本語   会議種別:口頭発表(一般)  

    researchmap

  • LSTM-PHV: prediction of human-virus protein-protein interactions by LSTM with word2vec

    Sho Tsukiyama, Md Mehedi Hasan, Satoshi Fujii, Hiroyuki Kurata

    2021年日本バイオインフォマティクス学会年会  2021年9月 

     詳細を見る

    開催年月日: 2021年9月

    記述言語:英語   会議種別:口頭発表(一般)  

    researchmap

  • 注意機構付き LSTM によるウイル スとヒトタンパク質間相互作用の予測

    築山 翔, Md. Mehedi Hasan, 藤井 聡, 倉田 博之

    情報処理学会 (バイオ情報学研究会)  2020年12月 

     詳細を見る

    開催年月日: 2020年12月

    記述言語:日本語   会議種別:口頭発表(一般)  

    researchmap

  • Detection of covert-speech-related potentials.

    Sho Tsukiyama, Toshimasa Yamazaki

    電子情報通信学会(MEとバイオサイバネティックス研究会)  2020年3月 

     詳細を見る

    開催年月日: 2020年3月

    記述言語:英語   会議種別:口頭発表(一般)  

    researchmap

  • Discriminability among Japanese vowels using early components in silent-speech-related potentials.

    Tsukiyama Sho, Toshimasa Yamazaki

    電子情報通信学会 (福祉情報工学研究会)  2019年10月 

     詳細を見る

    開催年月日: 2019年10月

    記述言語:英語   会議種別:口頭発表(一般)  

    researchmap

  • Segmentation of EEGs during silent speech based on statistical discriminability.

    Tsukiyama Sho, Sunaba Ken, Yamazaki Toshimasa

    The Society of Instrument and Control Engineers 2019  2019年9月 

     詳細を見る

    開催年月日: 2019年9月

    記述言語:英語   会議種別:ポスター発表  

    researchmap

▼全件表示

受賞

  • 情報処理学会バイオ情報学研究会 学生奨励賞

    2021年4月  

     詳細を見る

  • 情報処理学会バイオ情報学研究会第 65 回研究会 プレゼンテーション賞

    2020年9月  

     詳細を見る

共同研究・競争的資金等の研究課題

  • RNAの複数の安定立体構造と平衡分布を予測する革新的拡散モデルの開発

    研究課題/領域番号:25K24407  2025年7月 - 2027年3月

    日本学術振興会  科学研究費助成事業  研究活動スタート支援

    築山 翔

      詳細を見る

    配分額:2730000円 ( 直接経費:2100000円 、 間接経費:630000円 )

    researchmap

  • 革新的AIによるヒト-ウイルスタンパク質間相互作用阻害剤の予測

    研究課題/領域番号:22KJ2495  2023年3月 - 2025年3月

    日本学術振興会  科学研究費助成事業  特別研究員奨励費

    築山 翔

      詳細を見る

    配分額:3400000円 ( 直接経費:3400000円 )

    本研究ではウイルスの感染メカニズムの解明と抗ウイルス薬の開発に関連した研究を促進するために、アミノ酸配列からヒト-ウイルスタンパク質間相互作用と薬物-標的タンパク質間相互作用を高い精度で予測できるような手法の開発に取り組んでいる。当該年度はタンパク質のアミノ酸配列に加えて、タンパク質の構造情報を用いた予測を検討した。タンパク質の構造情報は、相互作用の予測を考える上で、原子の幾何学的な特徴を与える非常に重要な情報であるが、実験的にタンパク質の構造情報を同定することは容易ではなく、そのような情報を用いた相互作用の予測では構造が既知のタンパク質に適用が限られる。そのため、当初の研究計画ではアミノ酸配列のみを用いた相互作用予測手法の開発を検討していた。その一方で、近年、深層学習を用いることで、アミノ酸配列からタンパク質の構造情報を高い精度で予測することのできる手法が開発されている。そこで、このような手法を用いることで、アミノ酸配列からタンパク質の構造情報を予測し、アミノ酸配列とタンパク質の構造情報の両方を用いた相互作用の予測手法を検討した。このような研究の遂行において、研究活動の一部を海外の研究機関にて行うことで、生体分子の構造情報を用いた予測手法(前処理、特徴量の生成、モデル構築等)についての最先端の技術を調査し、機械学習や深層学習手法に加えて、シミュレーション技術などのアプローチを検討した。

    researchmap