-
Simplified Relative Citation Ratio for Static Paper Ranking: UFMG/LATIN at WSDM Cup 2016
Authors:
Sabir Ribas,
Alberto Ueda,
Rodrygo L. T. Santos,
Berthier Ribeiro-Neto,
Nivio Ziviani
Abstract:
Static rankings of papers play a key role in the academic search setting. Many features are commonly used in the literature to produce such rankings, some examples are citation-based metrics, distinct applications of PageRank, among others. More recently, learning to rank techniques have been successfully applied to combine sets of features producing effective results. In this work, we propose the…
▽ More
Static rankings of papers play a key role in the academic search setting. Many features are commonly used in the literature to produce such rankings, some examples are citation-based metrics, distinct applications of PageRank, among others. More recently, learning to rank techniques have been successfully applied to combine sets of features producing effective results. In this work, we propose the metric S-RCR, which is a simplified version of a metric called Relative Citation Ratio --- both based on the idea of a co-citation network. When compared to the classical version, our simplification S-RCR leads to improved efficiency with a reasonable effectiveness. We use S-RCR to rank over 120 million papers in the Microsoft Academic Graph dataset. By using this single feature, which has no parameters and does not need to be tuned, our team was able to reach the 3rd position in the first phase of the WSDM Cup 2016.
△ Less
Submitted 3 March, 2016;
originally announced March 2016.
-
P-score: A Publication-based Metric for Academic Productivity
Authors:
Sabir Ribas,
Berthier Ribeiro-Neto,
Edmundo de Souza e Silva,
Alberto Ueda,
Nivio Ziviani
Abstract:
In this work we propose a metric to assess academic productivity based on publication outputs. We are interested in knowing how well a research group in an area of knowledge is doing relatively to a pre-selected set of reference groups, where each group is composed by academics or researchers. To assess academic productivity we propose a new metric, which we call P-score. Our metric P-score assign…
▽ More
In this work we propose a metric to assess academic productivity based on publication outputs. We are interested in knowing how well a research group in an area of knowledge is doing relatively to a pre-selected set of reference groups, where each group is composed by academics or researchers. To assess academic productivity we propose a new metric, which we call P-score. Our metric P-score assigns weights to venues using only the publication patterns of selected reference groups. This implies that P-score does not depend on citation data and thus, that it is simpler to compute particularly in contexts in which citation data is not easily available. Also, preliminary experiments suggest that P-score preserves strong correlation with citation-based metrics.
△ Less
Submitted 11 March, 2015;
originally announced March 2015.
-
R-Score: Reputation-based Scoring of Research Groups
Authors:
Sabir Ribas,
Berthier Ribeiro-Neto,
Edmundo de Souza e Silva,
Nivio Ziviani
Abstract:
To manage the problem of having a higher demand for resources than availability of funds, research funding agencies usually rank the major research groups in their area of knowledge. This ranking relies on a careful analysis of the research groups in terms of their size, number of PhDs graduated, research results and their impact, among other variables. While research results are not the only vari…
▽ More
To manage the problem of having a higher demand for resources than availability of funds, research funding agencies usually rank the major research groups in their area of knowledge. This ranking relies on a careful analysis of the research groups in terms of their size, number of PhDs graduated, research results and their impact, among other variables. While research results are not the only variable to consider, they are frequently given special attention because of the notoriety they confer to the researchers and the programs they are affiliated with. In here we introduce a new metric for quantifying publication output, called R-Score for reputation-based score, which can be used in support to the ranking of research groups or programs. The novelty is that the metric depends solely on the listings of the publications of the members of a group, with no dependency on citation counts. R-Score has some interesting properties: (a) it does not require access to the contents of published material, (b) it can be curated to produce highly accurate results, and (c) it can be naturally used to compare publication output of research groups (e.g., graduate programs) inside a same country, geographical area, or across the world. An experiment comparing the publication output of 25 CS graduate programs from Brazil suggests that R-Score can be quite useful for providing early insights into the publication patterns of the various research groups one wants to compare.
△ Less
Submitted 27 August, 2013; v1 submitted 23 August, 2013;
originally announced August 2013.
-
Capacity Planning for Vertical Search Engines
Authors:
Claudine Badue,
Jussara Almeida,
Virgilio Almeida,
Ricardo Baeza-Yates,
Berthier Ribeiro-Neto,
Artur Ziviani,
Nivio Ziviani
Abstract:
Vertical search engines focus on specific slices of content, such as the Web of a single country or the document collection of a large corporation. Despite this, like general open web search engines, they are expensive to maintain, expensive to operate, and hard to design. Because of this, predicting the response time of a vertical search engine is usually done empirically through experimentation,…
▽ More
Vertical search engines focus on specific slices of content, such as the Web of a single country or the document collection of a large corporation. Despite this, like general open web search engines, they are expensive to maintain, expensive to operate, and hard to design. Because of this, predicting the response time of a vertical search engine is usually done empirically through experimentation, requiring a costly setup. An alternative is to develop a model of the search engine for predicting performance. However, this alternative is of interest only if its predictions are accurate. In this paper we propose a methodology for analyzing the performance of vertical search engines. Applying the proposed methodology, we present a capacity planning model based on a queueing network for search engines with a scale typically suitable for the needs of large corporations. The model is simple and yet reasonably accurate and, in contrast to previous work, considers the imbalance in query service times among homogeneous index servers. We discuss how we tune up the model and how we apply it to predict the impact on the query response time when parameters such as CPU and disk capacities are changed. This allows a manager of a vertical search engine to determine a priori whether a new configuration of the system might keep the query response under specified performance constraints.
△ Less
Submitted 25 June, 2010;
originally announced June 2010.