<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">Current Bioinformatics</journal-id><journal-title-group><journal-title xml:lang="en">Current Bioinformatics</journal-title><trans-title-group xml:lang="ru"><trans-title>Current Bioinformatics</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1574-8936</issn><issn publication-format="electronic">2212-392X</issn><publisher><publisher-name xml:lang="en">Bentham Science</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">643753</article-id><article-id pub-id-type="doi">10.2174/0115748936276861240109045208</article-id><article-categories><subj-group subj-group-type="toc-heading"><subject>Life Sciences</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">Genotype and Phenotype Association Analysis Based on Multi-omics Statistical Data</article-title></title-group><contrib-group><contrib contrib-type="author"><name><surname>Guo</surname><given-names>Xinpeng</given-names></name><email>info@benthamscience.net</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><name><surname>Song</surname><given-names>Yafei</given-names></name><email>info@benthamscience.net</email><xref ref-type="aff" rid="aff2"/></contrib><contrib contrib-type="author"><name><surname>Xu</surname><given-names>Dongyan</given-names></name><email>info@benthamscience.net</email><xref ref-type="aff" rid="aff2"/></contrib><contrib contrib-type="author"><name><surname>Jin</surname><given-names>Xueping</given-names></name><email>info@benthamscience.net</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><name><surname>Shang</surname><given-names>Xuequn</given-names></name><email>info@benthamscience.net</email><xref ref-type="aff" rid="aff3"/></contrib></contrib-group><aff id="aff1"><institution>School of Air and Missile Defense, Air Force Engineering University</institution></aff><aff id="aff2"><institution>Department of Basic Sciences, Air Force Engineering University</institution></aff><aff id="aff3"><institution>School of Computer Science and Engineering, Northwestern Polytechnical University</institution></aff><pub-date date-type="pub" iso-8601-date="2024-10-01" publication-format="electronic"><day>01</day><month>10</month><year>2024</year></pub-date><volume>19</volume><issue>10</issue><issue-title xml:lang="ru"/><fpage>933</fpage><lpage>942</lpage><history><date date-type="received" iso-8601-date="2025-01-07"><day>07</day><month>01</month><year>2025</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2024, Bentham Science Publishers</copyright-statement><copyright-year>2024</copyright-year><copyright-holder xml:lang="en">Bentham Science Publishers</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/></permissions><self-uri xlink:href="https://journals.eco-vector.com/1574-8936/article/view/643753">https://journals.eco-vector.com/1574-8936/article/view/643753</self-uri><abstract xml:lang="en"><p id="idm46041443811120">Background:When using clinical data for multi-omics analysis, there are issues such as the insufficient number of omics data types and relatively small sample size due to the protection of patients' privacy, the requirements of data management by various institutions, and the relatively large number of features of each omics data. This paper describes the analysis of multi-omics pathway relationships using statistical data in the absence of clinical data.</p><p id="idm46041443815120">Methods:We proposed a novel approach to exploit easily accessible statistics in public databases. This approach introduces phenotypic associations that are not included in the clinical data and uses these data to build a three-layer heterogeneous network. To simplify the analysis, we decomposed the three-layer network into double two-layer networks to predict the weights of the inter-layer associations. By adding a hyperparameter β, the weights of the two layers of the network were merged, and then k-fold cross-validation was used to evaluate the accuracy of this method. In calculating the weights of the two-layer networks, the RWR with fixed restart probability was combined with PBMDA and CIPHER to generate the PCRWR with biased weights and improved accuracy.</p><p id="idm46041443819088">Results:The area under the receiver operating characteristic curve was increased by approximately 7% in the case of the RWR with initial weights.</p><p id="idm46041443824144">Conclusion:Multi-omics statistical data were used to establish genotype and phenotype correlation networks for analysis, which was similar to the effect of clinical multi-omics analysis.</p></abstract><kwd-group xml:lang="en"><kwd>Genotype</kwd><kwd>phenotype</kwd><kwd>multi-omics statistical data</kwd><kwd>hyperparameter β</kwd><kwd>PBMDA</kwd><kwd>CIPHER.</kwd></kwd-group></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Guo X, Song Y, Liu S, Gao M, Qi Y, Shang X. Linking genotype to phenotype in multi-omics data of small sample. BMC Genomics 2021; 22(1): 537. doi: 10.1186/s12864-021-07867-w PMID: 34256701</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>Guo X, Han J, Song Y, Yin Z, Liu S, Shang X. Using expression quantitative trait loci data and graph-embedded neural networks to uncover genotypephenotype interactions. Front Genet 2022; 13: 921775. doi: 10.3389/fgene.2022.921775 PMID: 36046233</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Guo Y, Liu S, Li Z, Shang X. BCDForest: A boosting cascade deep forest model towards the classification of cancer subtypes based on gene expression data. BMC Bioinformatics 2018; 19(S5) (Suppl. 5): 118. doi: 10.1186/s12859-018-2095-4 PMID: 29671390</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>Guo X, Lu Y, Yin Z, Shang X. IPMM: Cancer subtype clustering model based on multiomics data and pathway and motif information. Cham: Springer International Publishing 2020; pp. 560-8.</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Fiscon G, Conte F, Farina L, Paci P. SAveRUNNER: A network-based algorithm for drug repurposing and its application to COVID-19. PLOS Comput Biol 2021; 17(2): e1008686. doi: 10.1371/journal.pcbi.1008686 PMID: 33544720</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>van Driel MA, Bruggeman J, Vriend G, Brunner HG, Leunissen JAM. A text-mining analysis of the human phenome. Eur J Hum Genet 2006; 14(5): 535-42. doi: 10.1038/sj.ejhg.5201585 PMID: 16493445</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>Kim Y, Park JH, Cho YR. Network-based approaches for disease-gene association prediction using protein-protein interaction networks. Int J Mol Sci 2022; 23(13): 7411. doi: 10.3390/ijms23137411 PMID: 35806415</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>Wu X, Jiang R, Zhang MQ, Li S. Network-based global inference of human disease genes. Mol Syst Biol 2008; 4(1): 189. doi: 10.1038/msb.2008.27 PMID: 18463613</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>Gilad Y, Rifkin SA, Pritchard JK. Revealing the architecture of gene regulation: the promise of eQTL studies. Trends Genet 2008; 24(8): 408-15. doi: 10.1016/j.tig.2008.06.001 PMID: 18597885</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>Schadt EE, Lamb J, Yang X, et al. An integrative genomics approach to infer causal associations between gene expression and disease. Nat Genet 2005; 37(7): 710-7. doi: 10.1038/ng1589 PMID: 15965475</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>Zhu Z, Zhang F, Hu H, et al. Integration of summary data from GWAS and eQTL studies predicts complex trait gene targets. Nat Genet 2016; 48(5): 481-7. doi: 10.1038/ng.3538 PMID: 27019110</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>Roytman M, Kichaev G, Gusev A, Pasaniuc B. Methods for fine-mapping with chromatin and expression data. PLoS Genet 2018; 14(2): e1007240. doi: 10.1371/journal.pgen.1007240 PMID: 29481575</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>Köhler S, Gargano M, Matentzoglu N, et al. The human phenotype ontology in 2021. Nucleic Acids Res 2021; 49(D1): D1207-17. doi: 10.1093/nar/gkaa1043 PMID: 33264411</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>Murtagh F, Contreras P. Algorithms for hierarchical clustering: An overview. Wiley Interdiscip Rev Data Min Knowl Discov 2012; 2(1): 86-97. doi: 10.1002/widm.53</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>Havens TC, Bezdek JC, Leckie C, Hall LO, Palaniswami M. Fuzzy c-means algorithms for very large data. IEEE Trans Fuzzy Syst 2012; 20(6): 1130-46. doi: 10.1109/TFUZZ.2012.2201485</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>Kohonen T. The self-organizing map. Neurocomputing 1998; 21(1-3): 1-6. doi: 10.1016/S0925-2312(98)00030-7</mixed-citation></ref><ref id="B17"><label>17.</label><mixed-citation>Wu FX. Genetic weighted k-means algorithm for clustering large-scale gene expression data. BMC Bioinformatics 2008; 9(S6) (Suppl. 6): S12. doi: 10.1186/1471-2105-9-S6-S12 PMID: 18541047</mixed-citation></ref><ref id="B18"><label>18.</label><mixed-citation>You ZH, Huang ZA, Zhu Z, et al. PBMDA: A novel and effective path-based computational model for miRNA-disease association prediction. PLOS Comput Biol 2017; 13(3): e1005455. doi: 10.1371/journal.pcbi.1005455 PMID: 28339468</mixed-citation></ref><ref id="B19"><label>19.</label><mixed-citation>Ba-alawi W, Soufan O, Essack M, Kalnis P, Bajic VB. DASPfind: new efficient method to predict drugtarget interactions. J Cheminform 2016; 8(1): 15. doi: 10.1186/s13321-016-0128-4 PMID: 26985240</mixed-citation></ref><ref id="B20"><label>20.</label><mixed-citation>Luo J, Long Y. NTSHMDA: Prediction of human microbe-disease association based on random walk by integrating network topological similarity. IEEE/ACM Trans Comput Biol Bioinform 2020; 17: 1341-51.</mixed-citation></ref><ref id="B21"><label>21.</label><mixed-citation>Köhler S, Bauer S, Horn D, Robinson PN. Walking the interactome for prioritization of candidate disease genes. Am J Hum Genet 2008; 82(4): 949-58. doi: 10.1016/j.ajhg.2008.02.013 PMID: 18371930</mixed-citation></ref><ref id="B22"><label>22.</label><mixed-citation>Li Y, Patra JC. Genome-wide inferring genephenotype relationship by walking on the heterogeneous network. Bioinformatics 2010; 26(9): 1219-24. doi: 10.1093/bioinformatics/btq108 PMID: 20215462</mixed-citation></ref><ref id="B23"><label>23.</label><mixed-citation>Chen X, Liu MX, Yan GY. RWRMDA: Predicting novel human microRNAdisease associations. Mol Biosyst 2012; 8(10): 2792-8. doi: 10.1039/c2mb25180a PMID: 22875290</mixed-citation></ref><ref id="B24"><label>24.</label><mixed-citation>Smedley D, Haider S, Durinck S, et al. The BioMart community portal: An innovative alternative to large, centralized data repositories. Nucleic Acids Res 2015; 43(W1): W589-98. doi: 10.1093/nar/gkv350 PMID: 25897122</mixed-citation></ref><ref id="B25"><label>25.</label><mixed-citation>Keshava Prasad TS, Goel R, Kandasamy K, et al. Human protein reference database-2009 update. Nucleic Acids Res 2009; 37(Database): D767-72. doi: 10.1093/nar/gkn892 PMID: 18988627</mixed-citation></ref><ref id="B26"><label>26.</label><mixed-citation>Mathivanan S, Ahmed M, Ahn NG, et al. Human Proteinpedia enables sharing of human protein data. Nat Biotechnol 2008; 26(2): 164-7. doi: 10.1038/nbt0208-164 PMID: 18259167</mixed-citation></ref><ref id="B27"><label>27.</label><mixed-citation>Piñero J, Bravo À, Queralt-Rosinach N, et al. DisGeNET: A comprehensive platform integrating information on human disease-associated genes and variants. Nucleic Acids Res 2017; 45(D1): D833-9. doi: 10.1093/nar/gkw943 PMID: 27924018</mixed-citation></ref><ref id="B28"><label>28.</label><mixed-citation>Peng J, Hui W, Li Q, et al. A learning-based framework for miRNA-disease association identification using neural networks. Bioinformatics 2019; 35(21): 4364-71. doi: 10.1093/bioinformatics/btz254 PMID: 30977780</mixed-citation></ref><ref id="B29"><label>29.</label><mixed-citation>Ramos EM, Hoffman D, Junkins HA, et al. Phenotypegenotype integrator (PheGenI): Synthesizing genome-wide association study (GWAS) data with existing genomic resources. Eur J Hum Genet 2014; 22(1): 144-7. doi: 10.1038/ejhg.2013.96 PMID: 23695286</mixed-citation></ref><ref id="B30"><label>30.</label><mixed-citation>Cornish AJ, David A, Sternberg MJE. PhenoRank: Reducing study bias in gene prioritization through simulation. Bioinformatics 2018; 34(12): 2087-95. doi: 10.1093/bioinformatics/bty028 PMID: 29360927</mixed-citation></ref><ref id="B31"><label>31.</label><mixed-citation>Zhang Y, Liu J, Liu X, et al. Prioritizing disease genes with an improved dual label propagation framework. BMC Bioinformatics 2018; 19(1): 47. doi: 10.1186/s12859-018-2040-6 PMID: 29422030</mixed-citation></ref><ref id="B32"><label>32.</label><mixed-citation>Yang K, Wang R, Liu G, et al. HerGePred: Heterogeneous network embedding representation for disease gene prediction. IEEE J Biomed Health Inform 2019; 23(4): 1805-15. doi: 10.1109/JBHI.2018.2870728 PMID: 31283472</mixed-citation></ref></ref-list></back></article>
