In the current biomedical data movement, numerous efforts have been made to convert and normalize a large number of traditional structured and unstructured data (e.g., EHRs, reports) to semi-structured data (e.g., RDF, OWL). With the increasing number of semi-structured data coming into the biomedical community, data integration and knowledge discovery from heterogeneous domains become important research problem. In the application level, detection of related concepts among medical ontologies is an important goal of life science research. It is more crucial to figure out how different concepts are related within a single ontology or across multiple ontologies by analysing predicates in different knowledge bases. However, the world today is one of information explosion, and it is extremely difficult for biomedical researchers to find existing or potential predicates to perform linking among cross domain concepts without any support from schema pattern analysis. Therefore, there is a need for a mechanism to do predicate oriented pattern analysis to partition heterogeneous ontologies into closer small topics and do query generation to discover cross domain knowledge from each topic. In this paper, we present such a model that predicates oriented pattern analysis based on their close relationship and generates a similarity matrix. Based on this similarity matrix, we apply an innovated unsupervised learning algorithm to partition large data sets into smaller and closer topics and generate meaningful queries to fully discover knowledge over a set of interlinked data sources. We have implemented a prototype system named BmQGen and evaluate the proposed model with colorectal surgical cohort from the Mayo Clinic.
Schmachtenberg, M., Bizer, C., Jentzsch, A. and Cyganiak, R. (2014) Linking Open Data Cloud Diagram 2014. http://lod-cloud.net/
Semantic Web Health Care and Life Sciences Interest Group. http://www.w3.org/2001/sw/hcls/
Garde, S., Knaup, P., Hovenga, E.J. and Heard, S. (2007) Towards Semantic Interoperability for Electronic Health Records—Domain Knowledge Governance for Open EHR Archetypes. Methods of Information in Medicine, 46, 332-343.
Shaw, M., et al. (2008) Generating Application Ontologies from Reference Ontologies. AMIA Annual Symposium Proceedings, 2008, 672-676.
Nekrutenko, A. and Taylor, J. (2012) Next-Generation Sequencing Data Interpretation: Enhancing Reproducibility and Accessibility. Nature Reviews Genetics, 13, 667-672. http://dx.doi.org/10.1038/nrg3305
Detwiler, L.T., Suciu, D. and Brinkley, J.F. (2008) Regular Paths in Sparql: Querying the Nci Thesaurus. AMIA Annual Symposium Proceedings, American Medical Informatics Association.
Sturn, A., Quackenbush, J. and Trajanoski, Z. (2002) Genesis: Cluster Analysis of Microarray Data. Bioinformatics, 18, 207-208. http://dx.doi.org/10.1093/bioinformatics/18.1.207
Dembélé, D. and Kastner, P. (2003) Fuzzy c-Means Method for Clustering Microarray Data. Bioinformatics, 19, 973-980. http://dx.doi.org/10.1093/bioinformatics/btg119
MedTQ: Dynamic Topic Discovery and Query Generation for Medical Ontologies. Manuscript Submitted for Publication, 10 January 2016. https://www.dropbox.com/s/9dcpidh5fjdxv1t/MedTQ.pdf?dl=0
Knowledge Discovery from Medical Ontologies in Cross Domains. Manuscript Submitted for Publication, 24 February 2016. https://www.dropbox.com/s/0dzhgqiqneokop0/MedKDD.pdf?dl=0
Shen, F.C. (2015) A Pervasive Framework for Real-Time Activity Patterns of Mobile Users. 2015 IEEE International Conference on Pervasive Computing and Communication Workshops, St. Louis, 23-27 March 2015, 248-250.
Shen, F.C., Liu, H.F., Sohn, S.H., Larson, D.W. and Lee, Y. (2015) BmQGen: Biomedical Query Generator for Knowledge Discovery. 2015 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Washington DC, 9-12 November 2015, 1092-1097.
MedTagger. http://ohnlp.org/index.php/MedTagger
Yates, A., Cafarella, M., Banko, M., Etzioni, O., Broadhead, M. and Soderland, S. (2007) Textrunner: Open Information Extraction on the Web. Proceedings of Human Language Technologies: The Annual Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, Rochester, 22-27 April 2007, 25-26. http://dx.doi.org/10.3115/1614164.1614177
Klyne, G. and Carroll, J.J. (2006) Resource Description Framework (RDF): Concepts and Abstract Syntax.
MedLex. http://www.medlex.com/
Liu, H.F., et al. (2014) Facilitating Post-Surgical Complication Detection through Sublanguage Analysis. Proceedings of the AMIA Joint Summits on Translational Science, 2014, 77-82.
Etzioni, O., Fader, A., Christensen, J., Soderland, S. and Mausam (2011) Open Information Extraction: The Second Generation. Proceedings of the 22nd International Joint Conference on Artificial Intelligence, 1, 3-10.
Fader, A., Soderland, S. and Etzioni, O. (2011) Identifying Relations for Open Information Extraction. Proceedings of the Conference on Empirical Methods in Natural Language Processing, Edinburgh, 27-29 July 2011, 1535-1545.
Chekol, M.W., Euzenat, J., Genevès, P. and Laya?da, N. (2011) PSPARQL Query Containment. The 13th International Symposium on Database Programming Languages, 29 August 2011, Seattle, 8 p.
Kochut, K.J. and Janik, M. (2007) SPARQLeR: Extended SPARQL for Semantic Association Discovery. In: Franconi, E., Kifer, M. and May, W., Eds., The Semantic Web: Research and Applications, Springer, Berlin, 145-159.
Bezdek, J.C., Ehrlich, R. and Full, W. (1984) FCM: The Fuzzy c-Means Clustering Algorithm. Computers & Geosciences, 10, 191-203. http://dx.doi.org/10.1016/0098-3004(84)90020-7
Eclipse Juno Integrated Development Environment. https://www.eclipse.org/juno/
McBride, B. (2001) Jena: Implementing the RDF Model and Syntax Specification. The 2nd International Workshop on the Semantic Web, Hongkong, 1 May 2001.
The R Project for Statistic. http://www.r-project.org/
Shannon, P., et al. (2003) Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. Genome Research, 13, 2498-2504. http://dx.doi.org/10.1101/gr.1239303
Hartigan, J.A. and Wong, M.A. (1979) Algorithm AS 136: A K-Means Clustering Algorithm. Journal of the Royal Statistical Society. Series C (Applied Statistics), 28, 100-108. http://dx.doi.org/10.2307/2346830
Ng, R.T. and Han, J.W. (2002) CLARANS: A Method for Clustering Objects for Spatial Data Mining. IEEE Transactions on Knowledge and Data Engineering, 14, 1003-1016. http://dx.doi.org/10.1109/TKDE.2002.1033770
Kaufman, L. and Rousseeuw, P.J. (1990) Partitioning around Medoids (Program PAM). In: Kaufman, L. and Rousseeuw, P.J., Eds., Finding Groups in Data: An Introduction to Cluster Analysis, John Wiley & Sons, Inc., Hoboken, 68-125.
Chen, B., Ding, Y. and Wild, D.J. (2012) Assessing Drug Target Association Using Semantic Linked Data. PLoS Computational Biology, 8, e1002574.
Doing-Harris, K., Livnat, Y. and Meystre, S. (2015) Automated Concept and Relationship Extraction for the Semi-Automated Ontology Management (SEAM) System. Journal of Biomedical Semantics, 6, 15.
Mate, S., et al. (2015) Ontology-Based Data Integration between Clinical and Research Systems. PloS ONE, 10, e0116656.
Willighagen, E.L., et al. (2013) The ChEMBL Database as Linked Open Data. Journal of Cheminformatics, 5, 23. http://dx.doi.org/10.1186/1758-2946-5-23
Belleau, F., Nolin, M.-A., Tourigny, N., Rigault, P. and Morissette, J. (2008) Bio2RDF: Towards a Mashup to Build Bioinformatics Knowledge Systems. Journal of Biomedical Informatics, 41, 706-716. http://dx.doi.org/10.1016/j.jbi.2008.03.004
Luciano, J.S., et al. (2011) The Translational Medicine Ontology and Knowledge Base: Driving Personalized Medicine by Bridging the Gap between Bench and Bedside. Journal of Biomedical Semantics, 2, S1.
Chen, B., et al. (2010) Chem2Bio2RDF: A Semantic Framework for Linking and Data Mining Chemogenomic and Systems Chemical Biology Data. BMC Bioinformatics, 11, 255.
Dumontier, M., et al. (2014) The Semanticscience Integrated Ontology (SIO) for Biomedical Research and Knowledge Discovery. Journal of Biomedical Semantics, 5, 14. http://dx.doi.org/10.1186/2041-1480-5-14
Croset, S., Hoehndorf, R. and Rebholz-Schuhmann, D. (2012) Integration of the Anatomical Therapeutic Chemical Classification System and DrugBank Using Owl and Text-Mining. Proceedings of the 4th Workshop of the GI Workgroup “Ontologies in Biomedicine and Life Sciences” (OBML), Dresden, 27-28 September 2012.
Chen, B., Ding, Y. and Wild, D.J. (2012) Improving Integrative Searching of Systems Chemical Biology Data Using Semantic Annotation. Journal of Cheminformatics, 4, 6. http://dx.doi.org/10.1186/1758-2946-4-6
Momtchev, V., Peychev, D., Primov, T. and Georgiev, G. (2009) Expanding the Pathway and Interaction Knowledge in Linked Life Data. Proceedings of the 8th International Semantic Web Challenge, Westfields Conference Center, 25-29 October 2009.
Samwald, M., et al. (2011) Linked Open Drug Data for Pharmaceutical Research and Development. Journal of Cheminformatics, 3, 19.
Hassanzadeh, O., Kementsietsidis, A., Lim, L., Miller, R.J. and Wang, M. (2009) LinkedCT: A Linked Data Space for Clinical Trials. arXiv:0908.0567.
Shi, C., Kong, X.N., Huang, Y., Yu, P.S. and Wu, B. (2014) Hetesim: A General Framework for Relevance Measure in Heterogeneous Networks. IEEE Transactions on Knowledge and Data Engineering, 26, 2479-2492. http://dx.doi.org/10.1109/TKDE.2013.2297920
Kong, X.N., Zhang, J.W. and Yu, P.S. (2013) Inferring Anchor Links across Multiple Heterogeneous Social Networks. Proceedings of the 22nd ACM international conference on Conference on Information & Knowledge Management, Burlingame, 27 October-1 November 2013, 179-188. http://dx.doi.org/10.1145/2505515.2505531
Garcia-Serna, R., Ursu, O., Oprea, T.I. and Mestres, J. (2010) IPHACE: Integrative Navigation in Pharmacological Space. Bioinformatics, 26, 985-986. http://dx.doi.org/10.1093/bioinformatics/btq061
Taboureau, O., et al. (2010) ChemProt: A Disease Chemical Biology Database. Nucleic Acids Research, 1-6.
Kuhn, M., Szklarczyk, D., Franceschini, A., von Mering, C., Jensen, L.J. and Bork, P. (2012) STITCH 3: Zooming in on Protein-Chemical Interactions. Nucleic Acids Research, 40, D876-D880. http://dx.doi.org/10.1093/nar/gkr1011
Oprea, T.I., et al. (2011) Associating Drugs, Targets and Clinical Outcomes into an Integrated Network Affords a New Platform for Computer-Aided Drug Repurposing. Molecular informatics, 30, 100-111.
Kinnings, S.L., Liu, N., Buchmeier, N., Tonge, P.J., Xie, L. and Bourne, P.E. (2009) Drug Discovery Using Chemical Systems Biology: Repositioning the Safe Medicine Comtan to Treat Multi-Drug and Extensively Drug Resistant Tuberculosis. PLoS Computational Biology, 5, e1000423.
Campillos, M., Kuhn, M., Gavin, A.-C., Jensen, L.J. and Bork, P. (2008) Drug Target Identification Using Side-Effect Similarity. Science, 321, 263-266. http://dx.doi.org/10.1126/science.1158140
Lamb, J., et al. (2006) The Connectivity Map: Using Gene-Expression Signatures to Connect Small Molecules, Genes, and Disease. Science, 313, 1929-1935.
Shi, B. and Weninger, T. (2015) Fact Checking in Large Knowledge Graphs-A Discriminative Predicate Path Mining Approach. arXiv:1510.05911.
Zhou, Y., Liu, L. and Buttler, D. (2015) Integrating Vertex-Centric Clustering with Edge-Centric Clustering for Meta Path Graph Analysis. Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Hilton, 10-13 August 2015, 1563-1572. http://dx.doi.org/10.1145/2783258.2783328
Van Leeuwen, M. (2014) Interactive Data Exploration Using Pattern Mining. In: Holzinger, A. and Jurisica, I., Eds., Interactive Knowledge Discovery and Data Mining in Biomedical Informatics, Springer, Berlin, 169-182. http://dx.doi.org/10.1007/978-3-662-43968-5_9
Gotz, D., Wang, F. and Perer, A. (2014) A Methodology for Interactive Mining and Visual Analysis of Clinical Event Patterns Using Electronic Health Record Data. Journal of Biomedical Informatics, 48, 148-159. http://dx.doi.org/10.1016/j.jbi.2014.01.007
K?lling, J., Langenk?mper, D., Abouna, S., Khan, M. and Nattkemper, T.W. (2012) WHIDE—A Web Tool for Visual Data Mining Colocation Patterns in Multivariate Bioimages. Bioinformatics, 28, 1143-1150. http://dx.doi.org/10.1093/bioinformatics/bts104
Huang, Z.X., Dong, W., Ji, L., Gan, C.X., Lu, X.D. and Duan, H.L. (2014) Discovery of Clinical Pathway Patterns from Event Logs Using Probabilistic Topic Models. Journal of Biomedical Informatics, 47, 39-57. http://dx.doi.org/10.1016/j.jbi.2013.09.003
Lasko, T.A., Denny, J.C. and Levy, M.A. (2013) Computational Phenotype Discovery Using Unsupervised Feature Learning over Noisy, Sparse, and Irregular Clinical Data. PloS ONE, 8, e66341.