Search | arXiv e-print repository

Well-Connected Communities in Real-World and Synthetic Networks

Authors: Minhyuk Park, Yasamin Tabatabaee, Vikram Ramavarapu, Baqiao Liu, Vidya Kamath Pailodi, Rajiv Ramachandran, Dmitriy Korobskiy, Fabio Ayres, George Chacko, Tandy Warnow

Abstract: Integral to the problem of detecting communities through graph clustering is the expectation that they are "well connected". In this respect, we examine five different community detection approaches optimizing different criteria: the Leiden algorithm optimizing the Constant Potts Model, the Leiden algorithm optimizing modularity, Iterative K-Core Clustering (IKC), Infomap, and Markov Clustering (M… ▽ More Integral to the problem of detecting communities through graph clustering is the expectation that they are "well connected". In this respect, we examine five different community detection approaches optimizing different criteria: the Leiden algorithm optimizing the Constant Potts Model, the Leiden algorithm optimizing modularity, Iterative K-Core Clustering (IKC), Infomap, and Markov Clustering (MCL). Surprisingly, all these methods produce, to varying extents, communities that fail even a mild requirement for well connectedness. To remediate clusters that are not well connected, we have developed the "Connectivity Modifier" (CM), which, at the cost of coverage, iteratively removes small edge cuts and re-clusters until all communities produced are well connected. Results from real-world and synthetic networks illustrate a tradeoff users make between well connected clusters and coverage, and raise questions about the "clusterability" of networks and models of community structure. △ Less

Submitted 14 August, 2023; v1 submitted 5 March, 2023; originally announced March 2023.

arXiv:2208.04842 [pdf, other]

doi 10.1162/qss_a_00227

AOC; Assembling Overlap** Communities

Authors: Akhil Jakatdar, Baqiao Liu, Tandy Warnow, George Chacko

Abstract: Through discovery of meso-scale structures, community detection methods contribute to the understanding of complex networks. Many community finding methods, however, rely on disjoint clustering techniques, in which node membership is restricted to one community or cluster. This strict requirement limits the ability to inclusively describe communities since some nodes may reasonably be assigned to… ▽ More Through discovery of meso-scale structures, community detection methods contribute to the understanding of complex networks. Many community finding methods, however, rely on disjoint clustering techniques, in which node membership is restricted to one community or cluster. This strict requirement limits the ability to inclusively describe communities since some nodes may reasonably be assigned to many communities. We have previously reported Iterative K-core Clustering (IKC), a scalable and modular pipeline that discovers disjoint research communities from the scientific literature. We now present Assembling Overlap** Clusters (AOC), a complementary meta-method for overlap** communities as an option that addresses the disjoint clustering problem. We present findings from the use of AOC on a network of over 13 million nodes that captures recent research in the very rapidly growing field of extracellular vesicles in biology. △ Less

Submitted 4 October, 2022; v1 submitted 5 August, 2022; originally announced August 2022.

Comments: This version submitted to Quantitative Science Studies

Journal ref: Quantitative Science Studies (2022)

arXiv:2111.07410 [pdf, other]

doi 10.1162/qss_a_00184

Center-Periphery Structure in Communities: Extracellular Vesicles

Authors: Eleanor Wedell, Minhyuk Park, Dmitriy Korobskiy, Tandy Warnow, George Chacko

Abstract: Clustering and community detection in networks are of broad interest and have been the subject of extensive research that spans several fields. We are interested in the relatively narrow question of detecting communities of scientific publications that are linked by citations. These publication communities can be used to identify scientists with shared interests who form communities of researchers… ▽ More Clustering and community detection in networks are of broad interest and have been the subject of extensive research that spans several fields. We are interested in the relatively narrow question of detecting communities of scientific publications that are linked by citations. These publication communities can be used to identify scientists with shared interests who form communities of researchers. Building on the well-known k-core algorithm, we have developed a modular pipeline to find publication communities. We compare our approach to communities discovered by the widely used Leiden algorithm for community finding. Using a quantitative and qualitative approach, we evaluate community finding results on a citation network consisting of over 14 million publications relevant to the field of extracellular vesicles. △ Less

Submitted 14 November, 2021; originally announced November 2021.

Journal ref: Quantitative Science Studies (2022)

arXiv:2007.14452 [pdf, other]

doi 10.1162/qss_a_00095

Finding Scientific Communities In Citation Graphs: Convergent Clustering

Authors: Shreya Chandrasekharan, Mariam Zaka, Stephen Gallo, Tandy Warnow, George Chacko

Abstract: Understanding the nature and organization of scientific communities is of broad interest. The `Invisible College' is a historical metaphor for one such type of community and the search for such `colleges' can be framed as the detection and analysis of small groups of scientists working on problems of common interests. Case studies have previously been conducted on individual communities with respe… ▽ More Understanding the nature and organization of scientific communities is of broad interest. The `Invisible College' is a historical metaphor for one such type of community and the search for such `colleges' can be framed as the detection and analysis of small groups of scientists working on problems of common interests. Case studies have previously been conducted on individual communities with respect to their scientific and social behavior. In this study, we introduce, a new and scalable community finding approach. Supplemented by expert assessment, we use the convergence of two different clustering methods to select article clusters generated from over two million articles from the field of immunology spanning an eleven year period with relevant cluster quality indicators for evaluation. Finally, we identify author communities defined by these clusters. A sample of the article clusters produced by this pipeline was reviewed by experts, and shows strong thematic relatedness, suggesting that the inferred author communities may represent valid communities of practice. These findings suggest that such convergent approaches may be useful in the future. △ Less

Submitted 28 July, 2020; originally announced July 2020.

Journal ref: Quantitative Science Studies (2021)

arXiv:2006.16737 [pdf, other]

doi 10.3389/frma.2020.577131

Delayed Recognition; the Co-citation Perspective

Authors: Wenxi Zhao, Dmitriy Korobskiy, George Chacko

Abstract: A Slee** Beauty is a publication that is apparently unrecognized for some period of time before experiencing sudden recognition by citation. Various reasons, including resistance to new ideas, have been attributed to such delayed recognition. We examine this phenomenon in the special case of co-citations, which represent new ideas generated through the combination of existing ones. Using relativ… ▽ More A Slee** Beauty is a publication that is apparently unrecognized for some period of time before experiencing sudden recognition by citation. Various reasons, including resistance to new ideas, have been attributed to such delayed recognition. We examine this phenomenon in the special case of co-citations, which represent new ideas generated through the combination of existing ones. Using relatively stringent selection criteria derived from the work of others, we analyze a very large dataset of over 940 million unique co-cited article pairs, and identified 1,196 cases of delayed co-citations. We further classify these 1,196 cases with respect to amplitude, rate of citation, and disciplinary origin and discuss alternative approaches towards identifying such instances. △ Less

Submitted 30 June, 2020; originally announced June 2020.

Journal ref: Frontiers in Research Metrics and Analytics (2021)

arXiv:2005.04793 [pdf, other]

doi 10.1162/qss_a_00075

Frequently Co-cited Publications: Features and Kinetics

Authors: Sitaram Devarakonda, James Bradley, Dmitriy Korobskiy, Tandy Warnow, George Chacko

Abstract: Co-citation measurements can reveal the extent to which a concept representing a novel combination of existing ideas evolves towards a specialty. The strength of co-citation is represented by its frequency, which accumulates over time. Of interest is whether underlying features associated with the strength of co-citation can be identified. We use the proximal citation network for a given pair of a… ▽ More Co-citation measurements can reveal the extent to which a concept representing a novel combination of existing ideas evolves towards a specialty. The strength of co-citation is represented by its frequency, which accumulates over time. Of interest is whether underlying features associated with the strength of co-citation can be identified. We use the proximal citation network for a given pair of articles (x, y) to compute theta, an a priori estimate of the probability of co-citation between x and y, prior to their first co-citation.Thus, low values for theta reflect pairs of articles for which co-citation is presumed less likely. We observe that co-citation frequencies are a composite of power-law and lognormal distributions, and that very high co-citation frequencies are more likely to be composed of pairs with low values of theta, reflecting the impact of a novel combination of ideas. Furthermore, we note that the occurrence of a direct citation between two members of a co-cited pair increases with co-citation frequency. Finally, we identify cases of frequently co-cited publications that accumulate co-citations after an extended period of dormancy. △ Less

Submitted 10 May, 2020; originally announced May 2020.

MSC Class: 01A85; 01A90 ACM Class: H.3.7

Journal ref: Quantitative Science Studies (2020)

arXiv:1912.10521 [pdf, other]

doi 10.1007/s11192-020-03624-0

Viewing Computer Science through Citation Analysis; Salton and Bergmark Redux

Authors: Sitaram Devarakonda, Dmitriy Korobskiy, Tandy Warnow, George Chacko

Abstract: Computer science has experienced dramatic growth and diversification over the last twenty years. Towards a current understanding of the structure of this discipline, we analyze a cohort of the computer science literature using the DBLP database. For insight on the features of this cohort and the relationship within its components, we constructed article level clusters based on either direct citati… ▽ More Computer science has experienced dramatic growth and diversification over the last twenty years. Towards a current understanding of the structure of this discipline, we analyze a cohort of the computer science literature using the DBLP database. For insight on the features of this cohort and the relationship within its components, we constructed article level clusters based on either direct citations or co-citations, and reconciled them to major and minor subject categories in the Scopus All Science Journal Classification (ASJC). We described complementary insights from clustering by direct citation and co-citation, and both point to the increase in computer science publications and their scope. Our analysis shows cross-category clusters, some that interact with external fields, such as the biological sciences, while others remain inward looking. △ Less

Submitted 22 December, 2019; originally announced December 2019.

MSC Class: K.2 ACM Class: K.2

arXiv:1911.08775 [pdf]

doi 10.1162/qss_a_00068

Do disruption index indicators measure what they propose to measure? The comparison of several indicator variants with assessments by peers

Authors: Lutz Bornmann, Sitaram Devarakonda, Alexander Tekles, George Chacko

Abstract: Recently, Wu, Wang, and Evans (2019) and Bu, Waltman, and Huang (2019) proposed a new family of indicators, which measure whether a scientific publication is disruptive to a field or tradition of research. Such disruptive influences are characterized by citations to a focal paper, but not its cited references. In this study, we are interested in the question of convergent validity, i.e., whether t… ▽ More Recently, Wu, Wang, and Evans (2019) and Bu, Waltman, and Huang (2019) proposed a new family of indicators, which measure whether a scientific publication is disruptive to a field or tradition of research. Such disruptive influences are characterized by citations to a focal paper, but not its cited references. In this study, we are interested in the question of convergent validity, i.e., whether these indicators of disruption are able to measure what they propose to measure ('disruptiveness'). We used external criteria of newness to examine convergent validity: in the post-publication peer review system of F1000Prime, experts assess papers whether the reported research fulfills these criteria (e.g., reports new findings). This study is based on 120,179 papers from F1000Prime published between 2000 and 2016. In the first part of the study we discuss the indicators. Based on the insights from the discussion, we propose alternate variants of disruption indicators. In the second part, we investigate the convergent validity of the indicators and the (possibly) improved variants. Although the results of a factor analysis show that the different variants measure similar dimensions, the results of regression analyses reveal that one variant (DI5) performs slightly better than the others. △ Less

Submitted 20 November, 2019; originally announced November 2019.

arXiv:1909.08738 [pdf, other]

doi 10.1162/qss_a_00007

Co-citations in context: disciplinary heterogeneity is relevant

Authors: James Bradley, Sitaram Devarakonda, Avon Davey, Dmitriy Korobskiy, Siyu Liu, Djamil Lakhdar-Hamina, Tandy Warnow, George Chacko

Abstract: Citation analysis of the scientific literature has been used to study and define disciplinary boundaries, to trace the dissemination of knowledge, and to estimate impact. Co-citation, the frequency with which pairs of publications are cited, provides insight into how documents relate to each other and across fields. Co-citation analysis has been used to characterize combinations of prior work as c… ▽ More Citation analysis of the scientific literature has been used to study and define disciplinary boundaries, to trace the dissemination of knowledge, and to estimate impact. Co-citation, the frequency with which pairs of publications are cited, provides insight into how documents relate to each other and across fields. Co-citation analysis has been used to characterize combinations of prior work as conventional or innovative and to derive features of highly cited publications. Given the organization of science into disciplines, a key question is the sensitivity of such analyses to frame of reference. Our study examines this question using semantically-themed citation networks. We observe that trends reported to be true across the scientific literature do not hold for focused citation networks, and we conclude that inferring novelty using co-citation analysis and random graph models benefits from disciplinary context. △ Less

Submitted 18 September, 2019; originally announced September 2019.

Journal ref: Quantitative Science Studies Oct 11, 2019

Showing 1–9 of 9 results for author: Chacko, G