location: Current position: Home >> Scientific Research >> Paper Publications

Prediction of LncRNA by Using Muitiple Feature Information Fusion and Feature Selection Technique

Hits:

Indexed by:会议论文

Date of Publication:2018-01-01

Included Journals:CPCI-S

Volume:10955

Page Number:318-329

Key Words:Ensemble feature selection; Maximum correlation minimum redundancy; Pseudo nucleotides features; Classification; LncRNA

Abstract:Recent genomic studies suggest that long non-coding RNAs (lncRNAs) play an important role in regulation of plant growth. Therefore, it is important to find more plant lncRNAs and predict their functions. This paper presents an improved maximum correlation minimum redundancy method for lncRNAs recognition. Sequence feature, secondary structural feature and functional feature such as pseudo-nucleotides feature which is based on the physical and chemical properties between dimers dinucleotide of related RNA have been extracted. Then, using maximum correlation minimum redundancy method to integrate a variety of feature selection methods such as Pearson correlation coefficient, information gain, relief algorithm and random forest for feature selection. Based on the selected superior feature subset, the classification model is established by SVM. Experimental results on Arabidopsis sequence dataset show that pseudo-nucleotides feature reflects information of different RNA sequences and the classification model constructed according to the proposed method can be more accurate than other methods on identification of plant lncRNAs.

Pre One:Overexpression of MiR482c in Tomato Induces Enhanced Susceptibility to Late Blight

Next One:Tomato lncRNA23468 functions as a competing endogenous RNA to modulate NBS-LRR genes by decoying miR482b in the tomato-Phytophthora infestans interaction