-
An Optimized Framework for Processing Large-scale Polysomnographic Data Incorporating Expert Human Oversight
Authors:
Benedikt Holm,
Gabriel Jouan,
Emil Hardarson,
Sigríður Sigurðardottir,
Kenan Hoelke,
Conor Murphy,
Erna Sif Arnardóttir,
María Óskarsdóttir,
Anna Sigríður Islind
Abstract:
Polysomnographic recordings are essential for diagnosing many sleep disorders, yet their detailed analysis presents considerable challenges. With the rise of machine learning methodologies, researchers have created various algorithms to automatically score and extract clinically relevant features from polysomnography, but less research has been devoted to how exactly the algorithms should be incor…
▽ More
Polysomnographic recordings are essential for diagnosing many sleep disorders, yet their detailed analysis presents considerable challenges. With the rise of machine learning methodologies, researchers have created various algorithms to automatically score and extract clinically relevant features from polysomnography, but less research has been devoted to how exactly the algorithms should be incorporated into the workflow of sleep technologists. This paper presents a sophisticated data collection platform developed under the Sleep Revolution project, to harness polysomnographic data from multiple European centers. A tripartite platform is presented: a user-friendly web platform for uploading three-night polysomnographic recordings, a dedicated splitter that segments these into individual one-night recordings, and an advanced processor that enhances the one-night polysomnography with contemporary automatic scoring algorithms. The platform is evaluated using real-life data and human scorers, whereby scoring time, accuracy and trust are quantified. Additionally, the scorers were interviewed about their trust in the platform, along with the impact of its integration into their workflow.
△ Less
Submitted 2 April, 2024;
originally announced April 2024.
-
Attention-based Dynamic Multilayer Graph Neural Networks for Loan Default Prediction
Authors:
Sahab Zandi,
Kamesh Korangi,
María Óskarsdóttir,
Christophe Mues,
Cristián Bravo
Abstract:
Whereas traditional credit scoring tends to employ only individual borrower- or loan-level predictors, it has been acknowledged for some time that connections between borrowers may result in default risk propagating over a network. In this paper, we present a model for credit risk assessment leveraging a dynamic multilayer network built from a Graph Neural Network and a Recurrent Neural Network, e…
▽ More
Whereas traditional credit scoring tends to employ only individual borrower- or loan-level predictors, it has been acknowledged for some time that connections between borrowers may result in default risk propagating over a network. In this paper, we present a model for credit risk assessment leveraging a dynamic multilayer network built from a Graph Neural Network and a Recurrent Neural Network, each layer reflecting a different source of network connection. We test our methodology in a behavioural credit scoring context using a dataset provided by U.S. mortgage financier Freddie Mac, in which different types of connections arise from the geographical location of the borrower and their choice of mortgage provider. The proposed model considers both types of connections and the evolution of these connections over time. We enhance the model by using a custom attention mechanism that weights the different time snapshots according to their importance. After testing multiple configurations, a model with GAT, LSTM, and the attention mechanism provides the best results. Empirical results demonstrate that, when it comes to predicting probability of default for the borrowers, our proposed model brings both better results and novel insights for the analysis of the importance of connections and timestamps, compared to traditional methods.
△ Less
Submitted 24 June, 2024; v1 submitted 31 January, 2024;
originally announced February 2024.
-
INFLECT-DGNN: Influencer Prediction with Dynamic Graph Neural Networks
Authors:
Elena Tiukhova,
Emiliano Penaloza,
María Óskarsdóttir,
Bart Baesens,
Monique Snoeck,
Cristián Bravo
Abstract:
Leveraging network information for predictive modeling has become widespread in many domains. Within the realm of referral and targeted marketing, influencer detection stands out as an area that could greatly benefit from the incorporation of dynamic network representation due to the ongoing development of customer-brand relationships. To elaborate this idea, we introduce INFLECT-DGNN, a new frame…
▽ More
Leveraging network information for predictive modeling has become widespread in many domains. Within the realm of referral and targeted marketing, influencer detection stands out as an area that could greatly benefit from the incorporation of dynamic network representation due to the ongoing development of customer-brand relationships. To elaborate this idea, we introduce INFLECT-DGNN, a new framework for INFLuencer prEdiCTion with Dynamic Graph Neural Networks that combines Graph Neural Networks (GNN) and Recurrent Neural Networks (RNN) with weighted loss functions, the Synthetic Minority Oversampling TEchnique (SMOTE) adapted for graph data, and a carefully crafted rolling-window strategy. To evaluate predictive performance, we utilize a unique corporate data set with networks of three cities and derive a profit-driven evaluation methodology for influencer prediction. Our results show how using RNN to encode temporal attributes alongside GNNs significantly improves predictive performance. We compare the results of various models to demonstrate the importance of capturing graph representation, temporal dependencies, and using a profit-driven methodology for evaluation.
△ Less
Submitted 12 December, 2023; v1 submitted 16 July, 2023;
originally announced July 2023.
-
Influencer Detection with Dynamic Graph Neural Networks
Authors:
Elena Tiukhova,
Emiliano Penaloza,
María Óskarsdóttir,
Hernan Garcia,
Alejandro Correa Bahnsen,
Bart Baesens,
Monique Snoeck,
Cristián Bravo
Abstract:
Leveraging network information for prediction tasks has become a common practice in many domains. Being an important part of targeted marketing, influencer detection can potentially benefit from incorporating dynamic network representation. In this work, we investigate different dynamic Graph Neural Networks (GNNs) configurations for influencer detection and evaluate their prediction performance u…
▽ More
Leveraging network information for prediction tasks has become a common practice in many domains. Being an important part of targeted marketing, influencer detection can potentially benefit from incorporating dynamic network representation. In this work, we investigate different dynamic Graph Neural Networks (GNNs) configurations for influencer detection and evaluate their prediction performance using a unique corporate data set. We show that using deep multi-head attention in GNN and encoding temporal attributes significantly improves performance. Furthermore, our empirical evaluation illustrates that capturing neighborhood representation is more beneficial that using network centrality measures.
△ Less
Submitted 15 November, 2022;
originally announced November 2022.
-
Gaining Insights on Student Course Selection in Higher Education with Community Detection
Authors:
Erla Guðrún Sturludóttir,
Eydís Arnardóttir,
Gísli Hjálmtýsson,
María Óskarsdóttir
Abstract:
Gaining insight into course choices holds significant value for universities, especially those who aim for flexibility in their programs and wish to adapt quickly to changing demands of the job market. However, little emphasis has been put on utilizing the large amount of educational data to understand these course choices. Here, we use network analysis of the course selection of all students who…
▽ More
Gaining insight into course choices holds significant value for universities, especially those who aim for flexibility in their programs and wish to adapt quickly to changing demands of the job market. However, little emphasis has been put on utilizing the large amount of educational data to understand these course choices. Here, we use network analysis of the course selection of all students who enrolled in an undergraduate program in engineering, psychology, business or computer science at a Nordic university over a five year period. With these methods, we have explored student choices to identify their distinct fields of interest. This was done by applying community detection to a network of courses, where two courses were connected if a student had taken both. We compared our community detection results to actual major specializations within the computer science department and found strong similarities. To compliment this analysis, we also used directed networks to identify the "typical" student, by looking at students' general course choices by semester. We found that course choices diversify as programs progress, meaning that attempting to understand course choices by identifying a "typical" student gives less insight than understanding what characterizes course choice diversity. Analysis with our proposed methodology can be used to offer more tailored education, which in turn allows students to follow their interests and adapt to the ever-changing career market.
△ Less
Submitted 4 May, 2021;
originally announced May 2021.
-
Effects of the COVID-19 Pandemic on Learning and Teaching: a Case Study from Higher Education
Authors:
Nidia Guadalupe López Flores,
Anna Sigridur Islind,
María Óskarsdóttir
Abstract:
In December 2019, the first case of SARS-CoV-2 infection was identified in Wuhan, China. Since that day, COVID-19 has spread worldwide, affecting 153 million people. Education, as many other sectors, has managed to adapt to the requirements and barriers implied by the impossibility to teach students face-to-face as it was done before. Yet, little is known about the implications of emergency remote…
▽ More
In December 2019, the first case of SARS-CoV-2 infection was identified in Wuhan, China. Since that day, COVID-19 has spread worldwide, affecting 153 million people. Education, as many other sectors, has managed to adapt to the requirements and barriers implied by the impossibility to teach students face-to-face as it was done before. Yet, little is known about the implications of emergency remote teaching (ERT) during the pandemic. This study describes and analyzes the impact of the pandemic on the study patterns of higher education students. The analysis was performed by the integration of three main components: (1) interaction with the learning management system (LMS), (2) Assignment submission rate, and (3) Teachers' perspective. Several variables were created to analyze the study patterns, clicks on different LMS components, usage during the day, week and part of the term, the time span of interaction with the LMS, and grade categories. The results showed significant differences in study patterns depending on the year of study, and the variables reflecting the effect of teachers' changes in the course structure are identified. This study outlines the first insights of higher education's new normality, providing important implications for supporting teachers in creating academic material that adequately addresses students' particular needs depending on their year of study, changes in study pattern, and distribution of time and activity through the term.
△ Less
Submitted 4 May, 2021;
originally announced May 2021.
-
Multilayer Network Analysis for Improved Credit Risk Prediction
Authors:
María Óskarsdóttir,
Cristián Bravo
Abstract:
We present a multilayer network model for credit risk assessment. Our model accounts for multiple connections between borrowers (such as their geographic location and their economic activity) and allows for explicitly modelling the interaction between connected borrowers. We develop a multilayer personalized PageRank algorithm that allows quantifying the strength of the default exposure of any bor…
▽ More
We present a multilayer network model for credit risk assessment. Our model accounts for multiple connections between borrowers (such as their geographic location and their economic activity) and allows for explicitly modelling the interaction between connected borrowers. We develop a multilayer personalized PageRank algorithm that allows quantifying the strength of the default exposure of any borrower in the network. We test our methodology in an agricultural lending framework, where it has been suspected for a long time default correlates between borrowers when they are subject to the same structural risks. Our results show there are significant predictive gains just by including centrality multilayer network information in the model, and these gains are increased by more complex information such as the multilayer PageRank variables. The results suggest default risk is highest when an individual is connected to many defaulters, but this risk is mitigated by the size of the neighbourhood of the individual, showing both default risk and financial stability propagate throughout the network.
△ Less
Submitted 26 July, 2021; v1 submitted 19 October, 2020;
originally announced October 2020.
-
Social network analytics for supervised fraud detection in insurance
Authors:
María Óskarsdóttir,
Waqas Ahmed,
Katrien Antonio,
Bart Baesens,
Rémi Dendievel,
Tom Donas,
Tom Reynkens
Abstract:
Insurance fraud occurs when policyholders file claims that are exaggerated or based on intentional damages. This contribution develops a fraud detection strategy by extracting insightful information from the social network of a claim. First, we construct a network by linking claims with all their involved parties, including the policyholders, brokers, experts, and garages. Next, we establish fraud…
▽ More
Insurance fraud occurs when policyholders file claims that are exaggerated or based on intentional damages. This contribution develops a fraud detection strategy by extracting insightful information from the social network of a claim. First, we construct a network by linking claims with all their involved parties, including the policyholders, brokers, experts, and garages. Next, we establish fraud as a social phenomenon in the network and use the BiRank algorithm with a fraud specific query vector to compute a fraud score for each claim. From the network, we extract features related to the fraud scores as well as the claims' neighborhood structure. Finally, we combine these network features with the claim-specific features and build a supervised model with fraud in motor insurance as the target variable. Although we build a model for only motor insurance, the network includes claims from all available lines of business. Our results show that models with features derived from the network perform well when detecting fraud and even outperform the models using only the classical claim-specific features. Combining network and claim-specific features further improves the performance of supervised learning models to detect fraud. The resulting model flags highly suspicions claims that need to be further investigated. Our approach provides a guided and intelligent selection of claims and contributes to a more effective fraud investigation process.
△ Less
Submitted 15 September, 2020;
originally announced September 2020.
-
Changes in mobility patterns in Europe during the COVID-19 pandemic: Novel insights using open source data
Authors:
Anna Sigridur Islind,
María Óskarsdóttir,
Harpa Steingrímsdóttir
Abstract:
The COVID-19 pandemic has changed the way we act, interact and move around in the world. The pandemic triggered a worldwide health crisis that has been tackled using a variety of strategies across Europe. Whereas some countries have taken strict measures, others have avoided lock-downs altogether. In this paper, we report on findings obtained by combining data from different publicly available sou…
▽ More
The COVID-19 pandemic has changed the way we act, interact and move around in the world. The pandemic triggered a worldwide health crisis that has been tackled using a variety of strategies across Europe. Whereas some countries have taken strict measures, others have avoided lock-downs altogether. In this paper, we report on findings obtained by combining data from different publicly available sources in order to shed light on the changes in mobility patterns in Europe during the pandemic. Using that data, we show that mobility patterns have changed in different counties depending on the strategies they adopted during the pandemic. Our data shows that the majority of European citizens walked less during the lock-downs, and that, even though flights were less frequent, driving increased drastically. In this paper, we focus on data for a number of countries, for which we have also developed a dashboard that can be used by other researchers for further analyses. Our work shows the importance of granularity in open source data and how such data can be used to shed light on the effects of the pandemic.
△ Less
Submitted 24 August, 2020;
originally announced August 2020.
-
Evolution of Credit Risk Using a Personalized Pagerank Algorithm for Multilayer Networks
Authors:
Cristián Bravo,
María Óskarsdóttir
Abstract:
In this paper we present a novel algorithm to study the evolution of credit risk across complex multilayer networks. Pagerank-like algorithms allow for the propagation of an influence variable across single networks, and allow quantifying the risk single entities (nodes) are subject to given the connection they have to other nodes in the network. Multilayer networks, on the other hand, are network…
▽ More
In this paper we present a novel algorithm to study the evolution of credit risk across complex multilayer networks. Pagerank-like algorithms allow for the propagation of an influence variable across single networks, and allow quantifying the risk single entities (nodes) are subject to given the connection they have to other nodes in the network. Multilayer networks, on the other hand, are networks where subset of nodes can be associated to a unique set (layer), and where edges connect elements either intra or inter networks. Our personalized PageRank algorithm for multilayer networks allows for quantifying how credit risk evolves across time and propagates through these networks. By using bipartite networks in each layer, we can quantify the risk of various components, not only the loans. We test our method in an agricultural lending dataset, and our results show how default risk is a challenging phenomenon that propagates and evolves through the network across time.
△ Less
Submitted 10 August, 2020; v1 submitted 25 May, 2020;
originally announced May 2020.
-
The Value of Big Data for Credit Scoring: Enhancing Financial Inclusion using Mobile Phone Data and Social Network Analytics
Authors:
María Óskarsdóttir,
Cristián Bravo,
Carlos Sarraute,
Jan Vanthienen,
Bart Baesens
Abstract:
Credit scoring is without a doubt one of the oldest applications of analytics. In recent years, a multitude of sophisticated classification techniques have been developed to improve the statistical performance of credit scoring models. Instead of focusing on the techniques themselves, this paper leverages alternative data sources to enhance both statistical and economic model performance. The stud…
▽ More
Credit scoring is without a doubt one of the oldest applications of analytics. In recent years, a multitude of sophisticated classification techniques have been developed to improve the statistical performance of credit scoring models. Instead of focusing on the techniques themselves, this paper leverages alternative data sources to enhance both statistical and economic model performance. The study demonstrates how including call networks, in the context of positive credit information, as a new Big Data source has added value in terms of profit by applying a profit measure and profit-based feature selection. A unique combination of datasets, including call-detail records, credit and debit account information of customers is used to create scorecards for credit card applicants. Call-detail records are used to build call networks and advanced social network analytics techniques are applied to propagate influence from prior defaulters throughout the network to produce influence scores. The results show that combining call-detail records with traditional data in credit scoring models significantly increases their performance when measured in AUC. In terms of profit, the best model is the one built with only calling behavior features. In addition, the calling behavior features are the most predictive in other models, both in terms of statistical and economic performance. The results have an impact in terms of ethical use of call-detail records, regulatory implications, financial inclusion, as well as data sharing and privacy.
△ Less
Submitted 23 February, 2020;
originally announced February 2020.
-
Credit Scoring for Good: Enhancing Financial Inclusion with Smartphone-Based Microlending
Authors:
María Óskarsdóttir,
Cristián Bravo,
Carlos Sarraute,
Bart Baesens,
Jan Vanthienen
Abstract:
Globally, two billion people and more than half of the poorest adults do not use formal financial services. Consequently, there is increased emphasis on develo** financial technology that can facilitate access to financial products for the unbanked. In this regard, smartphone-based microlending has emerged as a potential solution to enhance financial inclusion.
We propose a methodology to impr…
▽ More
Globally, two billion people and more than half of the poorest adults do not use formal financial services. Consequently, there is increased emphasis on develo** financial technology that can facilitate access to financial products for the unbanked. In this regard, smartphone-based microlending has emerged as a potential solution to enhance financial inclusion.
We propose a methodology to improve the predictive performance of credit scoring models used by these applications. Our approach is composed of several steps, where we mostly focus on engineering appropriate features from the user data. Thereby, we construct pseudo-social networks to identify similar people and combine complex network analysis with representation learning. Subsequently we build credit scoring models using advanced machine learning techniques with the goal of obtaining the most accurate credit scores, while also taking into consideration ethical and privacy regulations to avoid unfair discrimination. A successful deployment of our proposed methodology could improve the performance of microlending smartphone applications and help enhance financial wellbeing worldwide.
△ Less
Submitted 29 January, 2020;
originally announced January 2020.
-
Social Network Analytics for Churn Prediction in Telco: Model Building, Evaluation and Network Architecture
Authors:
María Óskarsdóttir,
Cristián Bravo,
Wouter Verbeke,
Carlos Sarraute,
Bart Baesens,
Jan Vanthienen
Abstract:
Social network analytics methods are being used in the telecommunication industry to predict customer churn with great success. In particular it has been shown that relational learners adapted to this specific problem enhance the performance of predictive models.
In the current study we benchmark different strategies for constructing a relational learner by applying them to a total of eight dist…
▽ More
Social network analytics methods are being used in the telecommunication industry to predict customer churn with great success. In particular it has been shown that relational learners adapted to this specific problem enhance the performance of predictive models.
In the current study we benchmark different strategies for constructing a relational learner by applying them to a total of eight distinct call-detail record datasets, originating from telecommunication organizations across the world. We statistically evaluate the effect of relational classifiers and collective inference methods on the predictive power of relational learners, as well as the performance of models where relational learners are combined with traditional methods of predicting customer churn in the telecommunication industry.
Finally we investigate the effect of network construction on model performance; our findings imply that the definition of edges and weights in the network does have an impact on the results of the predictive models. As a result of the study, the best configuration is a non-relational learner enriched with network variables, without collective inference, using binary weights and undirected networks. In addition, we provide guidelines on how to apply social networks analytics for churn prediction in the telecommunication industry in an optimal way, ranging from network architecture to model building and evaluation.
△ Less
Submitted 18 January, 2020;
originally announced January 2020.
-
A Comparative Study of Social Network Classifiers for Predicting Churn in the Telecommunication Industry
Authors:
Maria Óskarsdóttir,
Cristián Bravo,
Wouter Verbeke,
Carlos Sarraute,
Bart Baesens,
Jan Vanthienen
Abstract:
Relational learning in networked data has been shown to be effective in a number of studies. Relational learners, composed of relational classifiers and collective inference methods, enable the inference of nodes in a network given the existence and strength of links to other nodes. These methods have been adapted to predict customer churn in telecommunication companies showing that incorporating…
▽ More
Relational learning in networked data has been shown to be effective in a number of studies. Relational learners, composed of relational classifiers and collective inference methods, enable the inference of nodes in a network given the existence and strength of links to other nodes. These methods have been adapted to predict customer churn in telecommunication companies showing that incorporating them may give more accurate predictions. In this research, the performance of a variety of relational learners is compared by applying them to a number of CDR datasets originating from the telecommunication industry, with the goal to rank them as a whole and investigate the effects of relational classifiers and collective inference methods separately. Our results show that collective inference methods do not improve the performance of relational classifiers and the best performing relational classifier is the network-only link-based classifier, which builds a logistic model using link-based measures for the nodes in the network.
△ Less
Submitted 18 January, 2020;
originally announced January 2020.