-
Current and future directions in network biology
Authors:
Marinka Zitnik,
Michelle M. Li,
Aydin Wells,
Kimberly Glass,
Deisy Morselli Gysi,
Arjun Krishnan,
T. M. Murali,
Predrag Radivojac,
Sushmita Roy,
Anaïs Baudot,
Serdar Bozdag,
Danny Z. Chen,
Lenore Cowen,
Kapil Devkota,
Anthony Gitter,
Sara Gosline,
Pengfei Gu,
Pietro H. Guzzi,
Heng Huang,
Meng Jiang,
Ziynet Nesibe Kesimoglu,
Mehmet Koyuturk,
Jian Ma,
Alexander R. Pico,
Nataša Pržulj
, et al. (12 additional authors not shown)
Abstract:
Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These challenges stem from various fa…
▽ More
Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These challenges stem from various factors, notably the growing complexity and volume of data together with the increased diversity of data types describing different tiers of biological organization. We discuss prevailing research directions in network biology and highlight areas of inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Following the overview of recent breakthroughs across these five areas, we offer a perspective on the future directions of network biology. Additionally, we offer insights into scientific communities, educational initiatives, and the importance of fostering diversity within the field. This paper establishes a roadmap for an immediate and long-term vision for network biology.
△ Less
Submitted 11 June, 2024; v1 submitted 15 September, 2023;
originally announced September 2023.
-
Adverse Health Correlates of Intimate Partner Violence against Older Women: Mining Electronic Health Records
Authors:
Serhan Yilmaz,
Erkan Gunay,
Da Hee Lee,
Kathleen Whiting,
Kristen Silver,
Mehmet Koyuturk,
Gunnur Karakurt
Abstract:
Intimate partner violence (IPV) is often studied as a problem that predominantly affects younger women. However, studies show that older women are also frequently victims of abuse even though the physical effects of abuse are harder to detect. In this study, we mined the electronic health records (EHR) available through IBM Explorys to identify health correlates of IPV that are specific to older w…
▽ More
Intimate partner violence (IPV) is often studied as a problem that predominantly affects younger women. However, studies show that older women are also frequently victims of abuse even though the physical effects of abuse are harder to detect. In this study, we mined the electronic health records (EHR) available through IBM Explorys to identify health correlates of IPV that are specific to older women. Our analyses suggested that diagnostic terms that are co-morbid with IPV in older women are dominated by substance abuse and associated toxicities. When we considered differential co-morbidity, i.e., terms that are significantly more associated with IPV in older women compared to younger women, we identified terms spanning mental health issues, musculoskeletal issues, neoplasms, and disorders of various organ systems including skin, ears, nose and throat. Our findings provide pointers for further investigation in understanding the health effects of IPV among older women, as well as potential markers that can be used for screening IPV.
△ Less
Submitted 24 March, 2022;
originally announced March 2022.
-
Random Walks with Variable Restarts for Negative-Example-Informed Label Propagation
Authors:
Sean Maxwell,
Mehmet Koyuturk
Abstract:
Label propagation is frequently encountered in machine learning and data mining applications on graphs, either as a standalone problem or as part of node classification. Many label propagation algorithms utilize random walks (or network propagation), which provide limited ability to take into account negatively-labeled nodes (i.e., nodes that are known to be not associated with the label of intere…
▽ More
Label propagation is frequently encountered in machine learning and data mining applications on graphs, either as a standalone problem or as part of node classification. Many label propagation algorithms utilize random walks (or network propagation), which provide limited ability to take into account negatively-labeled nodes (i.e., nodes that are known to be not associated with the label of interest). Specialized algorithms to incorporate negatively labeled samples generally focus on learning or readjusting the edge weights to drive walks away from negatively-labeled nodes and toward positively-labeled nodes. This approach has several disadvantages, as it increases the number of parameters to be learned, and does not necessarily drive the walk away from regions of the network that are rich in negatively-labeled nodes.
We reformulate random walk with restarts and network propagation to enable "variable restarts", that is the increased likelihood of restarting at a positively-labeled node when a negatively-labeled node is encountered. Based on this reformulation, we develop CusTaRd, an algorithm that effectively combines variable restart probabilities and edge re-weighting to avoid negatively-labeled nodes. In addition to allowing variable restarts, CusTaRd samples negatively-labeled nodes from neighbors of positively-labeled nodes to better characterize the difference between positively and negatively labeled nodes. To assess the performance of CusTaRd, we perform comprehensive experiments on four network datasets commonly used in benchmarking label propagation and node classification algorithms. Our results show that CusTaRd consistently outperforms competing algorithms that learn/readjust edge weights, and sampling of negatives from the close neighborhood of positives further improves predictive accuracy.
△ Less
Submitted 13 October, 2021;
originally announced October 2021.
-
Expanding Label Sets for Graph Convolutional Networks
Authors:
Mustafa Coskun,
Burcu Bakir Gungor,
Mehmet Koyuturk
Abstract:
In recent years, Graph Convolutional Networks (GCNs) and their variants have been widely utilized in learning tasks that involve graphs. These tasks include recommendation systems, node classification, among many others. In node classification problem, the input is a graph in which the edges represent the association between pairs of nodes, multi-dimensional feature vectors are associated with the…
▽ More
In recent years, Graph Convolutional Networks (GCNs) and their variants have been widely utilized in learning tasks that involve graphs. These tasks include recommendation systems, node classification, among many others. In node classification problem, the input is a graph in which the edges represent the association between pairs of nodes, multi-dimensional feature vectors are associated with the nodes, and some of the nodes in the graph have known labels. The objective is to predict the labels of the nodes that are not labeled, using the nodes features, in conjunction with graph topology. While GCNs have been successfully applied to this problem, the caveats that they inherit from traditional deep learning models pose significant challenges to broad utilization of GCNs in node classification. One such caveat is that training a GCN requires a large number of labeled training instances, which is often not the case in realistic settings. To remedy this requirement, state-of-the-art methods leverage network diffusion-based approaches to propagate labels across the network before training GCNs. However, these approaches ignore the tendency of the network diffusion methods in biasing proximity with centrality, resulting in the propagation of labels to the nodes that are well-connected in the graph. To address this problem, here we present an alternate approach to extrapolating node labels in GCNs in the following three steps: (i) clustering of the network to identify communities, (ii) use of network diffusion algorithms to quantify the proximity of each node to the communities, thereby obtaining a low-dimensional topological profile for each node, (iii) comparing these topological profiles to identify nodes that are most similar to the labeled nodes.
△ Less
Submitted 18 December, 2019;
originally announced December 2019.
-
Fast Computation of Katz Index for Efficient Processing of Link Prediction Queries
Authors:
Mustafa Coskun,
Abdelkader Baggag,
Mehmet Koyuturk
Abstract:
Network proximity computations are among the most common operations in various data mining applications, including link prediction and collaborative filtering. A common measure of network proximity is Katz index, which has been shown to be among the best-performing path-based link prediction algorithms. With the emergence of very large network databases, such proximity computations become an impor…
▽ More
Network proximity computations are among the most common operations in various data mining applications, including link prediction and collaborative filtering. A common measure of network proximity is Katz index, which has been shown to be among the best-performing path-based link prediction algorithms. With the emergence of very large network databases, such proximity computations become an important part of query processing in these databases. Consequently, significant effort has been devoted to develo** algorithms for efficient computation of Katz index between a given pair of nodes or between a query node and every other node in the network. Here, we present LRC-Katz, an algorithm based on indexing and low-rank correction to accelerate Katz index-based network proximity queries. Using a variety of very large real-world networks, we show that LRC-Katz outperforms the fastest existing method, Conjugate Gradient, for a wide range of parameter values. We also show that this acceleration in the computation of Katz index can be used to drastically improve the efficiency of processing link prediction queries in very large networks. Motivated by this observation, we propose a new link prediction algorithm that exploits modularity of networks that are encountered in practical applications. Our experimental results on the link prediction problem show that our modularity based algorithm significantly outperforms the state-of-the-art link prediction Katz method.
△ Less
Submitted 13 December, 2019;
originally announced December 2019.