H. Wang
Please Note
34 records found
1
are weighted according to their contribution to the spreading process. Six hyperedge
properties were evaluated as predictors of backbone weights on six hypergraph datasets
for two infection rates (β) using Pearson and Kendall correlation. The properties
Hyperdegree, Degree, and Closeness Centrality generally exhibited negative correlations with the backbone weights, whereas Betweenness Centrality and Shortest Detour
showed positive correlations, particularly for higher infection rates. Neighbourhood
Coefficient showed strong dependence on both the infection rate and the dataset structure. In general, global properties performed better for larger values of β, while local
properties showed stronger predictive power on lower values of β. No single property
consistently achieved the strongest correlations across all datasets. Instead, the predictive power of hyperedge properties appears to depend on the infection rate and the
underlying structure of the hypergraph.
...
are weighted according to their contribution to the spreading process. Six hyperedge
properties were evaluated as predictors of backbone weights on six hypergraph datasets
for two infection rates (β) using Pearson and Kendall correlation. The properties
Hyperdegree, Degree, and Closeness Centrality generally exhibited negative correlations with the backbone weights, whereas Betweenness Centrality and Shortest Detour
showed positive correlations, particularly for higher infection rates. Neighbourhood
Coefficient showed strong dependence on both the infection rate and the dataset structure. In general, global properties performed better for larger values of β, while local
properties showed stronger predictive power on lower values of β. No single property
consistently achieved the strongest correlations across all datasets. Instead, the predictive power of hyperedge properties appears to depend on the infection rate and the
underlying structure of the hypergraph.
The main contribution of this paper is the proposal of a novel strategy for adding hyperlinks to the hypergraph to increase information spread. This novel method, NIPHD, takes into account the approximate probability of a node being infected and the number of hyperlinks to which it is connected. This strategy vastly outperforms others in 6 real-world datasets. With it, we observed increases of up to 10% in the number of infected nodes when doubling the number of edges connecting 3 nodes. This corresponds to a 5% increase compared to random addition, although this varies per hypergraph. ...
The main contribution of this paper is the proposal of a novel strategy for adding hyperlinks to the hypergraph to increase information spread. This novel method, NIPHD, takes into account the approximate probability of a node being infected and the number of hyperlinks to which it is connected. This strategy vastly outperforms others in 6 real-world datasets. With it, we observed increases of up to 10% in the number of infected nodes when doubling the number of edges connecting 3 nodes. This corresponds to a 5% increase compared to random addition, although this varies per hypergraph.
Vulnerability of Information Transport on Hypergraphs to Hyperlink Removal
An Empirical Study of Shortest-Path Routing Robustness
This paper presents a comparative evaluation of five hyperlink-removal strategies based on concepts previously proposed in the network science literature. By comparing their effects on shortest-path information transport, we assess which structural properties best indicate hyperlink importance and hypergraph vulnerability. The strategies are evaluated on six real-world datasets. To enable fair comparisons across structurally different hypergraphs, hyperlinks of different cardinalities (or orders) are analyzed separately, and strategy performance is compared after removing equivalent proportions of hyperlinks. Performance is assessed through the reduction in global efficiency, whereas the size of the largest connected component provides a complementary measure of structural fragmentation.
The results show that strategies incorporating global structural information consistently outperform local heuristics. In particular, Hyperlink Betweenness Centrality achieves the largest reduction in both global efficiency and connectivity across all provided datasets. The findings further demonstrate that hypergraph structure strongly influences vulnerability, with larger and more overlapping hyperlinks increasing robustness to targeted attacks. These results provide new insights into hyperlink importance and the factors governing resilience in higher-order networks. ...
This paper presents a comparative evaluation of five hyperlink-removal strategies based on concepts previously proposed in the network science literature. By comparing their effects on shortest-path information transport, we assess which structural properties best indicate hyperlink importance and hypergraph vulnerability. The strategies are evaluated on six real-world datasets. To enable fair comparisons across structurally different hypergraphs, hyperlinks of different cardinalities (or orders) are analyzed separately, and strategy performance is compared after removing equivalent proportions of hyperlinks. Performance is assessed through the reduction in global efficiency, whereas the size of the largest connected component provides a complementary measure of structural fragmentation.
The results show that strategies incorporating global structural information consistently outperform local heuristics. In particular, Hyperlink Betweenness Centrality achieves the largest reduction in both global efficiency and connectivity across all provided datasets. The findings further demonstrate that hypergraph structure strongly influences vulnerability, with larger and more overlapping hyperlinks increasing robustness to targeted attacks. These results provide new insights into hyperlink importance and the factors governing resilience in higher-order networks.
Inhibiting Spread in Hypergraphs through Community-Aware and Process-Based Node Removal
A comparative study under the SICP model
Rewiring Hypergraphs to Improve Information Propagation
The Impact of Degree Preserving Rewiring on the SICP Model
Influence Maximization in Temporal Networks
Heuristic and Sketch Based Approaches
tion guarantee, the highest mean relative performance of all methods tested across all datasets, and scales to networks of hundreds of thousands of nodes. The other heuristic methods proposed in this work trade some influence for lower computational complexity, all having lower influence than Temporal IMM. ...
tion guarantee, the highest mean relative performance of all methods tested across all datasets, and scales to networks of hundreds of thousands of nodes. The other heuristic methods proposed in this work trade some influence for lower computational complexity, all having lower influence than Temporal IMM.
Decentralization in DeFi Lending: A Network Perspective
A Multi-Layered Analysis of Governance-Active Users
We benchmark four methods - Baseline, SD, SCD, and an extended SCD* - across a range of physical (face-to-face) and virtual (online communication) contact networks, each exhibiting unique structural and dynamic properties when aggregated into temporal weighted networks. Our results show that the SCD model achieves a 32.86% reduction in Mean Squared Error (MSE) over the Baseline, while SD yields a 19.31% improvement. Moving from SD to SCD confers an additional 15.87% decrease in MSE, underscoring the benefits of incorporating both temporal and structural information. Although SCD* introduces further complexity, it did not show any consistent improvements in predictive performance across datasets.
Additional evaluations using the Area Under the Precision-Recall Curve (AUPRC) highlight dataset-specific variability in capturing active links. Correlation analysis reveals that MSE scales with average link weight distributions, whereas AUPRC correlates strongly with the proportion of active links per network snapshot. These findings emphasize that incorporating decay and structural context, in an interpretable manner, significantly enhances predictive accuracy, although parameter tuning remains crucial for different network topologies and interaction patterns. ...
We benchmark four methods - Baseline, SD, SCD, and an extended SCD* - across a range of physical (face-to-face) and virtual (online communication) contact networks, each exhibiting unique structural and dynamic properties when aggregated into temporal weighted networks. Our results show that the SCD model achieves a 32.86% reduction in Mean Squared Error (MSE) over the Baseline, while SD yields a 19.31% improvement. Moving from SD to SCD confers an additional 15.87% decrease in MSE, underscoring the benefits of incorporating both temporal and structural information. Although SCD* introduces further complexity, it did not show any consistent improvements in predictive performance across datasets.
Additional evaluations using the Area Under the Precision-Recall Curve (AUPRC) highlight dataset-specific variability in capturing active links. Correlation analysis reveals that MSE scales with average link weight distributions, whereas AUPRC correlates strongly with the proportion of active links per network snapshot. These findings emphasize that incorporating decay and structural context, in an interpretable manner, significantly enhances predictive accuracy, although parameter tuning remains crucial for different network topologies and interaction patterns.
We introduce FIMH, an algorithm for fair influence maximization on hypergraphs. Operating under the Susceptible–Infected Contact Process (SICP) model, FIMH jointly optimizes total influence and fairness across communities using a structural influence estimation and a parameter-free utopia-distance selection criterion. Experiments on seven real-world hypergraph datasets demonstrate that FIMH has competitive influence performance to state-of-the-art methods while reducing inter-community disparity by 31% on average and up to 52%. Our results establish that fairness and influence are not competing objectives in hypergraph diffusion such that balanced information spread can be achieved without sacrificing reach.
...
We introduce FIMH, an algorithm for fair influence maximization on hypergraphs. Operating under the Susceptible–Infected Contact Process (SICP) model, FIMH jointly optimizes total influence and fairness across communities using a structural influence estimation and a parameter-free utopia-distance selection criterion. Experiments on seven real-world hypergraph datasets demonstrate that FIMH has competitive influence performance to state-of-the-art methods while reducing inter-community disparity by 31% on average and up to 52%. Our results establish that fairness and influence are not competing objectives in hypergraph diffusion such that balanced information spread can be achieved without sacrificing reach.
Spreading Processes on Networks
Roles of Nodes, Links, and Hyperlinks
We first explore how the network properties of a node can be used to predict the spreading influence of the node, defined as the average number of nodes that are ultimately infected when this node is the only seed node. Previous studies have shown that combining node properties derived from local and global topological information can better predict nodal influence than using a single metric. In Chapter 2, we investigate whether using relatively local information is sufficient for the prediction. To address this question, we define an iterative metric set by leveraging the iterative process used to derive classical nodal centralities like eigenvector centrality. The iterative metric set progressively incorporates more global information and is used as the feature set in a regression model to predict nodal spreading influence. The iterative metric set is then used as the feature set in a regression model to predict the spreading influence of a node. We find that the model using the iterative metric set that includes relatively local information achieves comparable prediction quality with the method that includes both local and global information, in various networks.
A spreading process can be mitigated by blocking social contacts, i.e., time-specific interactions. In Chapter 3, we investigate how the network properties of a contact are associated with the mitigation effect when the contact is blocked. We develop probabilistic contact blocking strategies, which remove contacts (temporal links) based on their properties in a temporal network, to mitigate the spread of a Susceptible-Infected-Recovered spreading process. The removal probability of a contact depends on a given centrality metric of the corresponding link in the time-aggregated network and the occurring time of the contact. We propose diverse link centrality metrics, and each centrality metric leads to a unique contact blocking strategy. Our results indicate that the spread of the epidemic is most effectively mitigated when contacts between node pairs that have fewer contacts and contacts that occur earlier in time are more likely to be removed.
The role of a link in a spreading process can also be reflected by the extent to which the link is used in the process. Many real-world systems may involve interactions among groups of more than two individuals and can therefore be represented as temporal higher-order networks. Chapter 4 explores the Susceptible-Infected threshold spreading process unfolding on temporal higher-order networks with two objectives: (1) to understand the contribution of each hyperlink to the spreading process, defined as the average number of nodes that are directly infected via the activation of the hyperlink starting from an arbitrary seed node, and (2) to investigate hyperlinks with what network properties tend to contribute more to the spreading process. This understanding is crucial for developing effective strategies to mitigate a spreading process. Given a temporal higher-order network, we propose to construct a weighted higher-order network, the so-called diffusion backbone, where the weight of each hyperlink denotes its contribution to the spreading process. We then systematically design centrality metrics for hyperlinks in a temporal higher-order network, where each centrality metric captures a specific property of the hyperlink within a temporal higher-order network and is used to estimate the ranking of hyperlinks by their weights in the backbone. We find and explain why certain centrality metrics can better estimate the contributions of hyperlinks under different parameters of the spreading process.
The last chapter reflects on the insights of this thesis and discusses possible future directions related to our research. ...
We first explore how the network properties of a node can be used to predict the spreading influence of the node, defined as the average number of nodes that are ultimately infected when this node is the only seed node. Previous studies have shown that combining node properties derived from local and global topological information can better predict nodal influence than using a single metric. In Chapter 2, we investigate whether using relatively local information is sufficient for the prediction. To address this question, we define an iterative metric set by leveraging the iterative process used to derive classical nodal centralities like eigenvector centrality. The iterative metric set progressively incorporates more global information and is used as the feature set in a regression model to predict nodal spreading influence. The iterative metric set is then used as the feature set in a regression model to predict the spreading influence of a node. We find that the model using the iterative metric set that includes relatively local information achieves comparable prediction quality with the method that includes both local and global information, in various networks.
A spreading process can be mitigated by blocking social contacts, i.e., time-specific interactions. In Chapter 3, we investigate how the network properties of a contact are associated with the mitigation effect when the contact is blocked. We develop probabilistic contact blocking strategies, which remove contacts (temporal links) based on their properties in a temporal network, to mitigate the spread of a Susceptible-Infected-Recovered spreading process. The removal probability of a contact depends on a given centrality metric of the corresponding link in the time-aggregated network and the occurring time of the contact. We propose diverse link centrality metrics, and each centrality metric leads to a unique contact blocking strategy. Our results indicate that the spread of the epidemic is most effectively mitigated when contacts between node pairs that have fewer contacts and contacts that occur earlier in time are more likely to be removed.
The role of a link in a spreading process can also be reflected by the extent to which the link is used in the process. Many real-world systems may involve interactions among groups of more than two individuals and can therefore be represented as temporal higher-order networks. Chapter 4 explores the Susceptible-Infected threshold spreading process unfolding on temporal higher-order networks with two objectives: (1) to understand the contribution of each hyperlink to the spreading process, defined as the average number of nodes that are directly infected via the activation of the hyperlink starting from an arbitrary seed node, and (2) to investigate hyperlinks with what network properties tend to contribute more to the spreading process. This understanding is crucial for developing effective strategies to mitigate a spreading process. Given a temporal higher-order network, we propose to construct a weighted higher-order network, the so-called diffusion backbone, where the weight of each hyperlink denotes its contribution to the spreading process. We then systematically design centrality metrics for hyperlinks in a temporal higher-order network, where each centrality metric captures a specific property of the hyperlink within a temporal higher-order network and is used to estimate the ranking of hyperlinks by their weights in the backbone. We find and explain why certain centrality metrics can better estimate the contributions of hyperlinks under different parameters of the spreading process.
The last chapter reflects on the insights of this thesis and discusses possible future directions related to our research.
Lacking precise knowledge of the behavior of the population of interest, we propose a robustness analysis. Here, we present a tool for generating synthetic datasets, based on well-established models for the movement of individuals. We used the tool to generate data for a range of behavioral properties, encompassing variations in both underlying movement and phone usage. We evaluated three existing methods using our synthetic data. The first is a discriminatory approach that learns typical movement patterns and phone usage from a reference dataset. The second approach uses a model of cell tower behavior, making minimal assumptions on user behavior by choosing pairs of registrations close in time. The third is a generic statistical method for comparing event data. Additionally, we present a fourth method that combines the latter two, as conceptually, they use different aspects of the data.
Our analysis reveals that the discriminatory method performs best in a baseline scenario but is most sensitive to behavioral deviations. The cell tower method shows the lowest baseline performance yet exhibits the strongest resilience to variations. The generic model appears intermediate in terms of performance and sensitivity. Given the importance of robustness in evaluating evidence, we recommend using the combined approach, which is both reliable and effective across our defined variations. ...
Lacking precise knowledge of the behavior of the population of interest, we propose a robustness analysis. Here, we present a tool for generating synthetic datasets, based on well-established models for the movement of individuals. We used the tool to generate data for a range of behavioral properties, encompassing variations in both underlying movement and phone usage. We evaluated three existing methods using our synthetic data. The first is a discriminatory approach that learns typical movement patterns and phone usage from a reference dataset. The second approach uses a model of cell tower behavior, making minimal assumptions on user behavior by choosing pairs of registrations close in time. The third is a generic statistical method for comparing event data. Additionally, we present a fourth method that combines the latter two, as conceptually, they use different aspects of the data.
Our analysis reveals that the discriminatory method performs best in a baseline scenario but is most sensitive to behavioral deviations. The cell tower method shows the lowest baseline performance yet exhibits the strongest resilience to variations. The generic model appears intermediate in terms of performance and sensitivity. Given the importance of robustness in evaluating evidence, we recommend using the combined approach, which is both reliable and effective across our defined variations.
This project aims to find intrinsic factors that influence RWA performance in WDM and propose novel RWA approaches with enhanced performance. Existing dynamic RWA methods are reviewed from the literature and simulated in a self-built performance evaluation model. As the availability of every edge at every wavelength is constantly changing, we can transform the WDM network into a multi-layer temporal network structure. In order to uncover the essential reasons for the differences between the performances of the different methods, we investigate the multi-layer temporal network with graph theoretic analysis to explore correlations between specific multi-layer metrics and RWA performance. A few single-layer network connectivity metrics are applied in multi-layer networks including the number of connected components, the size of the largest components, the spectral radius, the algebraic connectivity, the effective resistance, the sum of betweenness, and the number of reachable node pairs. The experimental results show that the maximum value of spectral radius and algebraic connectivity over all layers are the best 2 multi-layer metrics describing the performance of the RWA methods.
Building upon this analysis, four new routing methods are proposed based on the previous methods and the two best-adapted multi-layer graph metrics, including the Least Spectral Radius Deduction (LSRD), Least Algebraic Connectivity Deduction (LACD), Least Hopcount and Congestion Path (LHCP) and Congestion Weighted Shortest Path (CWSP) methods. All new methods are fitted in the evaluation model and it has been proven that the CWSP method has better performance compared to all other RWA methods based on its improvement of selected multi-layer graph metrics. ...
This project aims to find intrinsic factors that influence RWA performance in WDM and propose novel RWA approaches with enhanced performance. Existing dynamic RWA methods are reviewed from the literature and simulated in a self-built performance evaluation model. As the availability of every edge at every wavelength is constantly changing, we can transform the WDM network into a multi-layer temporal network structure. In order to uncover the essential reasons for the differences between the performances of the different methods, we investigate the multi-layer temporal network with graph theoretic analysis to explore correlations between specific multi-layer metrics and RWA performance. A few single-layer network connectivity metrics are applied in multi-layer networks including the number of connected components, the size of the largest components, the spectral radius, the algebraic connectivity, the effective resistance, the sum of betweenness, and the number of reachable node pairs. The experimental results show that the maximum value of spectral radius and algebraic connectivity over all layers are the best 2 multi-layer metrics describing the performance of the RWA methods.
Building upon this analysis, four new routing methods are proposed based on the previous methods and the two best-adapted multi-layer graph metrics, including the Least Spectral Radius Deduction (LSRD), Least Algebraic Connectivity Deduction (LACD), Least Hopcount and Congestion Path (LHCP) and Congestion Weighted Shortest Path (CWSP) methods. All new methods are fitted in the evaluation model and it has been proven that the CWSP method has better performance compared to all other RWA methods based on its improvement of selected multi-layer graph metrics.
Communication professionals at research institutes are tasked with connecting scientists and journalists. The recommender system supports this process by recommending scientist-journalist connections based on data from previous collaborations. A scientist collaboration network, a journalist collaboration network and a scientist-journalist collaboration network are combined into a multilayer network. A recommender system is designed based on centrality metrics in the scientist and journalist collaboration networks and distance metrics in the multilayer network. In contrast to traditional link prediction problems - which aim to predict what links are most likely to form in the network - the problem in this thesis is how to recommend the most likely link for a single node, i.e. the most likely scientist links for a given journalist or most likely journalist links for a given scientist. A novel evaluation method is created to evaluate the performance of the recommender system.
The development of this system is used as a vessel to research how participation in a digital development process affects the mental model of digital innovation. This research contributes to addressing the lack of understanding of how to develop a mental model that facilitates innovation in the context of digital transformation. Three themes were identified in their mental model change: The extent to which innovation requires involvement, the complexity of innovation processes and what outcomes can realistically be expected of a digital innovation process. The team went from a model of digital innovation as 'a mysterious black box' - something external, where they could hand in a list of requirements and walk away with a digital tool - to a 'super puppy' that can do remarkable things, but has to be trained and interacted with to get a desired effect. ...
Communication professionals at research institutes are tasked with connecting scientists and journalists. The recommender system supports this process by recommending scientist-journalist connections based on data from previous collaborations. A scientist collaboration network, a journalist collaboration network and a scientist-journalist collaboration network are combined into a multilayer network. A recommender system is designed based on centrality metrics in the scientist and journalist collaboration networks and distance metrics in the multilayer network. In contrast to traditional link prediction problems - which aim to predict what links are most likely to form in the network - the problem in this thesis is how to recommend the most likely link for a single node, i.e. the most likely scientist links for a given journalist or most likely journalist links for a given scientist. A novel evaluation method is created to evaluate the performance of the recommender system.
The development of this system is used as a vessel to research how participation in a digital development process affects the mental model of digital innovation. This research contributes to addressing the lack of understanding of how to develop a mental model that facilitates innovation in the context of digital transformation. Three themes were identified in their mental model change: The extent to which innovation requires involvement, the complexity of innovation processes and what outcomes can realistically be expected of a digital innovation process. The team went from a model of digital innovation as 'a mysterious black box' - something external, where they could hand in a list of requirements and walk away with a digital tool - to a 'super puppy' that can do remarkable things, but has to be trained and interacted with to get a desired effect.
Fairness Aware Influence
Study of relationship between node network properties and FAI in complex networks
The primary objective of this thesis is to study how network properties such as degree and community size, relate with FAI. Network properties are measured using centrality metrics, which are categorized into two types.
The first type, referred to as "simple" or classic centrality metrics, do not account for community information. The second type, known as community-aware centrality metrics, incorporate community information but were not originally designed for FAI. These serve as baselines for ranking nodes in terms of FAI. Thus, two new classes of metrics are designed specifically for FAI in the attempt to perform better than the baselines.
Six real world networks are employed to evaluate the metrics. Local centrality and Community-Hub-Bridge are found to be good baselines in their respective categories, and the newly proposed metrics surpass the existing ones at the epidemic threshold. Additionally, a discussion is presented to compare and analyze these metrics, considering their performance under varying infection rates using an SIR infection spreading model. ...
The primary objective of this thesis is to study how network properties such as degree and community size, relate with FAI. Network properties are measured using centrality metrics, which are categorized into two types.
The first type, referred to as "simple" or classic centrality metrics, do not account for community information. The second type, known as community-aware centrality metrics, incorporate community information but were not originally designed for FAI. These serve as baselines for ranking nodes in terms of FAI. Thus, two new classes of metrics are designed specifically for FAI in the attempt to perform better than the baselines.
Six real world networks are employed to evaluate the metrics. Local centrality and Community-Hub-Bridge are found to be good baselines in their respective categories, and the newly proposed metrics surpass the existing ones at the epidemic threshold. Additionally, a discussion is presented to compare and analyze these metrics, considering their performance under varying infection rates using an SIR infection spreading model.
In such a time of economic instability as caused by the coronavirus outbreak, it is very useful for a company to know in what financial state they are going to be such that they can actively take precautions, such as liquidating their assets or decreasing their expenses. The financial state of a company is often reflected using Key Performance Indicators, or KPIs for short. These KPIs include metrics like the revenue, cost, and cash flow of a company. The forecasting of these KPIs can help a company in informing in what financial state they are going to be and are usually done using historical data of the company. Whereas the decrease in economic activity of business partners of a company is not reflected in the historical KPI data of the company itself, it can be seen in a network of companies that indicates whether there exists a relationship between two companies by using data on monetary transactions between companies. For this reason, we think that enriching historical KPI data using node features extracted from a dynamic network of companies can help improve the quality of KPI predictions during a period of economic instability such as the COVID-19 pandemic.
This thesis answers the question of whether we can use utilize a dynamic network of SMEs to improve the quality of KPI predictions during the COVID-19 lockdown. To answer this question, we first focus on creating a dynamic network consisting of SMEs and the transactions between them out of unstandardized data by proposing a novel, lightweight entity resolution algorithm that is used to find a mapping between companies. The resulting network is analyzed, and we found that the effects of the coronavirus lockdown are visible in the network. Next, we examine several KPIs, such as the revenue or the cash flow of a company, and we found that we can also see the effects of the COVID-19 lockdown in several KPIs. Lastly, this thesis describes an analysis of whether node features can be used to improve the quality of the forecasting of several of these KPIs, where we found that node features such as the degree and clustering coefficient of a node can indeed help with improving KPI forecasting under certain conditions. ...
In such a time of economic instability as caused by the coronavirus outbreak, it is very useful for a company to know in what financial state they are going to be such that they can actively take precautions, such as liquidating their assets or decreasing their expenses. The financial state of a company is often reflected using Key Performance Indicators, or KPIs for short. These KPIs include metrics like the revenue, cost, and cash flow of a company. The forecasting of these KPIs can help a company in informing in what financial state they are going to be and are usually done using historical data of the company. Whereas the decrease in economic activity of business partners of a company is not reflected in the historical KPI data of the company itself, it can be seen in a network of companies that indicates whether there exists a relationship between two companies by using data on monetary transactions between companies. For this reason, we think that enriching historical KPI data using node features extracted from a dynamic network of companies can help improve the quality of KPI predictions during a period of economic instability such as the COVID-19 pandemic.
This thesis answers the question of whether we can use utilize a dynamic network of SMEs to improve the quality of KPI predictions during the COVID-19 lockdown. To answer this question, we first focus on creating a dynamic network consisting of SMEs and the transactions between them out of unstandardized data by proposing a novel, lightweight entity resolution algorithm that is used to find a mapping between companies. The resulting network is analyzed, and we found that the effects of the coronavirus lockdown are visible in the network. Next, we examine several KPIs, such as the revenue or the cash flow of a company, and we found that we can also see the effects of the COVID-19 lockdown in several KPIs. Lastly, this thesis describes an analysis of whether node features can be used to improve the quality of the forecasting of several of these KPIs, where we found that node features such as the degree and clustering coefficient of a node can indeed help with improving KPI forecasting under certain conditions.
Node Influence Prediction in Complex Networks
Towards network embedding based features
Recently, a limited number of studies have proposed methods on how to utilize classical network topology based features to predict the nodal influence. However, two main challenges still persist: (1) individual topology based features do not fully capture the information of a node and (2) it is tedious to obtain these features for nodes in large scale networks. As an alternative solution, this work aims to utilize network embedding based features instead, where feature vectors of the nodes are learned from the network topology. In this research we assume that the network topology and the nodal influence of a small subset of the nodes are known. We then proceed to show how to build and optimize a machine learning framework where only 10\% of the nodes are used as training data and which could even be applicable on large scale networks. Additionally, we also demonstrate why network embedding based features are applicable in the node influence prediction task.
The findings show that node pairs which are closer in proximity in the network, are also embedded closer in the embedding space (exhibiting a higher similarity). The performance evaluation of the predictive models illustrate that network embedding based features can compete with classical topological metrics, despite the disadvantage of their higher dimensionality. This is achieved by combining the embedding features with individual low cost topology features such as the degree. ...
Recently, a limited number of studies have proposed methods on how to utilize classical network topology based features to predict the nodal influence. However, two main challenges still persist: (1) individual topology based features do not fully capture the information of a node and (2) it is tedious to obtain these features for nodes in large scale networks. As an alternative solution, this work aims to utilize network embedding based features instead, where feature vectors of the nodes are learned from the network topology. In this research we assume that the network topology and the nodal influence of a small subset of the nodes are known. We then proceed to show how to build and optimize a machine learning framework where only 10\% of the nodes are used as training data and which could even be applicable on large scale networks. Additionally, we also demonstrate why network embedding based features are applicable in the node influence prediction task.
The findings show that node pairs which are closer in proximity in the network, are also embedded closer in the embedding space (exhibiting a higher similarity). The performance evaluation of the predictive models illustrate that network embedding based features can compete with classical topological metrics, despite the disadvantage of their higher dimensionality. This is achieved by combining the embedding features with individual low cost topology features such as the degree.
This thesis aims to predict the demand for the next week by applying white-box and black-box models. The problem is represented as a temporal weighted bipartite network prediction problem. Specifically, the goal is to predict the network structure at a time T+1 based on the bipartite network observed at time T-k+1, T-k+2, ..., T, where k is an integer and needs to be optimized such that the prediction error is minimized. By analyzing the data in its temporal dimension, using the autocorrelation and cross-correlation among different storage locations, floors and volumes, it was shown that the autocorrelation is high and the cross-correlation is low. This suggests that the temporal bipartite network is possibly predictable.
We have explored different state-of-the-art predictive techniques. Markov chain model, LSTM, and ConvLSTM have been selected because of their fundamental difference in the way they learn and predict. A performance comparison is given where the techniques have been applied on the storage data, and it shows that LSTM outperforms the Markov chain and ConvLSTM based on the following evaluation metrics: RMSE, MAE, and accuracy. According to our dataset, higher predictability was achieved when only the data of a single link was exploited. The Markov chain and the LSTM utilize the information of a single link to predict. On the contrary, the ConvLSTM utilizes the information of the entire network to predict. The low cross-correlation between the links explains why the LSTM outperforms the ConvLSTM. The ConvLSTM tries to capture spatio-temporal dependencies, while this, in general, does not contain much valuable predictive information. Thus, the model is introduced to more noise, making it harder to predict accurately. The LSTM also outperforms the Markov chain model, which is used as a baseline method. This proves that it is beneficial to use a complex deep learning model for this dataset to predict. However, the Markov chain performs comparable to the ConvLSTM, showing that a black-box model does not always outperform a white-box model. This emphasizes that the most suitable predictive algorithm depends on the statistical properties of the dataset.
The theoretical upper bound of the predictability of the network is computed. It is the upper bound that can be used to compare the realized performance to the maximum achievable prediction performance for any predictive algorithm. The difference between the performance of the best performing algorithm to our dataset, LSTM, and the theoretical upper bound is still large, indicating that there is still room for improvement.
...
This thesis aims to predict the demand for the next week by applying white-box and black-box models. The problem is represented as a temporal weighted bipartite network prediction problem. Specifically, the goal is to predict the network structure at a time T+1 based on the bipartite network observed at time T-k+1, T-k+2, ..., T, where k is an integer and needs to be optimized such that the prediction error is minimized. By analyzing the data in its temporal dimension, using the autocorrelation and cross-correlation among different storage locations, floors and volumes, it was shown that the autocorrelation is high and the cross-correlation is low. This suggests that the temporal bipartite network is possibly predictable.
We have explored different state-of-the-art predictive techniques. Markov chain model, LSTM, and ConvLSTM have been selected because of their fundamental difference in the way they learn and predict. A performance comparison is given where the techniques have been applied on the storage data, and it shows that LSTM outperforms the Markov chain and ConvLSTM based on the following evaluation metrics: RMSE, MAE, and accuracy. According to our dataset, higher predictability was achieved when only the data of a single link was exploited. The Markov chain and the LSTM utilize the information of a single link to predict. On the contrary, the ConvLSTM utilizes the information of the entire network to predict. The low cross-correlation between the links explains why the LSTM outperforms the ConvLSTM. The ConvLSTM tries to capture spatio-temporal dependencies, while this, in general, does not contain much valuable predictive information. Thus, the model is introduced to more noise, making it harder to predict accurately. The LSTM also outperforms the Markov chain model, which is used as a baseline method. This proves that it is beneficial to use a complex deep learning model for this dataset to predict. However, the Markov chain performs comparable to the ConvLSTM, showing that a black-box model does not always outperform a white-box model. This emphasizes that the most suitable predictive algorithm depends on the statistical properties of the dataset.
The theoretical upper bound of the predictability of the network is computed. It is the upper bound that can be used to compare the realized performance to the maximum achievable prediction performance for any predictive algorithm. The difference between the performance of the best performing algorithm to our dataset, LSTM, and the theoretical upper bound is still large, indicating that there is still room for improvement.