Search | arXiv e-print repository

HateDebias: On the Diversity and Variability of Hate Speech Debiasing

Authors: Nankai Lin, Hongyan Wu, Zhengming Chen, Zijian Li, Lianxi Wang, Shengyi Jiang, Dong Zhou, Aimin Yang

Abstract: Hate speech on social media is ubiquitous but urgently controlled. Without detecting and mitigating the biases brought by hate speech, different types of ethical problems. While a number of datasets have been proposed to address the problem of hate speech detection, these datasets seldom consider the diversity and variability of bias, making it far from real-world scenarios. To fill this gap, we p… ▽ More Hate speech on social media is ubiquitous but urgently controlled. Without detecting and mitigating the biases brought by hate speech, different types of ethical problems. While a number of datasets have been proposed to address the problem of hate speech detection, these datasets seldom consider the diversity and variability of bias, making it far from real-world scenarios. To fill this gap, we propose a benchmark, named HateDebias, to analyze the model ability of hate speech detection under continuous, changing environments. Specifically, to meet the diversity of biases, we collect existing hate speech detection datasets with different types of biases. To further meet the variability (i.e., the changing of bias attributes in datasets), we reorganize datasets to follow the continuous learning setting. We evaluate the detection accuracy of models trained on the datasets with a single type of bias with the performance on the HateDebias, where a significant performance drop is observed. To provide a potential direction for debiasing, we further propose a debiasing framework based on continuous learning and bias information regularization, as well as the memory replay strategies to ensure the debiasing ability of the model. Experiment results on the proposed benchmark show that the aforementioned method can improve several baselines with a distinguished margin, highlighting its effectiveness in real-world applications. △ Less

Submitted 7 June, 2024; originally announced June 2024.

arXiv:2406.04735 [pdf, other]

On the capability of high redshift kSZ measurement with galaxy surveys

Authors: Ziyang Chen, Pengjie Zhang

Abstract: The kSZ effect has been detected at z<1 using various techniques and data sets. The ongoing and upcoming spectroscopic galaxy surveys such as DESI and PFS will push the detection beyond z = 1, and therefore map the baryon distribution at high redshifts. Such detection can be achieved by both the kSZ stacking and tomography methods. While the two methods are theoretically equivalent, they differ si… ▽ More The kSZ effect has been detected at z<1 using various techniques and data sets. The ongoing and upcoming spectroscopic galaxy surveys such as DESI and PFS will push the detection beyond z = 1, and therefore map the baryon distribution at high redshifts. Such detection can be achieved by both the kSZ stacking and tomography methods. While the two methods are theoretically equivalent, they differ significantly in the probed physics and scales, and required data sets. Taking the combination of PFS and ACT as an example, we build mocks of kSZ and galaxies, quantify the kSZ detection S/N, and compare between the two methods. We segment the PFS galaxies into three redshift bins: 0.6 < z < 1.0, 1.0 < z < 1.6, and 1.6 < z < 2.4. For tomography method, our analysis reveals that the two higher redshift bins exhibit higher S/N, with values of 32 and 28, respectively, compared to the first redshift bin (S/N = 8). This is attributed to not only the increasing of electron density with redshifts, but also the larger survey volume and the reduced non-linearity, facilitating velocity reconstruction at higher redshifts. Therefore, the capability of the PFS survey to measure high redshift kSZ effect stands as a substantial advantage over other spectroscopic surveys at lower redshift. The S/N of kSZ stacking largely depends on the number of clusters/groups available from another photometric survey. But in general, its S/N is lower than that of kSZ tomography. Incorporating next-generation CMB surveys like CMB-S4, characterized by significantly reduced instrument noise and improved angular resolution, is expected to enhance tomographic detection by a factor of ten and stacking detection by five. This future high S/N detection holds the promise of not only providing precise constraints on the overall baryon abundance but also initiating a new insight into baryon distribution. △ Less

Submitted 7 June, 2024; originally announced June 2024.

Comments: 23 pages, 6 figures

arXiv:2406.04707 [pdf, ps, other]

Nonlinear Optimal Guidance with Constraints on Impact Time and Impact Angle

Authors: Fanchen Wu, Zheng Chen, Xueming Shao, Kun Wang

Abstract: This paper aims to address the nonlinear optimal guidance problem with impact-time and impact-angle constraints, which is fundamentally important for multiple pursuers to collaboratively achieve a target. Addressing such a guidance problem is equivalent to solving a nonlinear minimum-effort control problem in real time. To this end, the Pontryagain's maximum principle is employed to convert extrem… ▽ More This paper aims to address the nonlinear optimal guidance problem with impact-time and impact-angle constraints, which is fundamentally important for multiple pursuers to collaboratively achieve a target. Addressing such a guidance problem is equivalent to solving a nonlinear minimum-effort control problem in real time. To this end, the Pontryagain's maximum principle is employed to convert extremal trajectories as the solutions of a parameterized differential system. The geometric property for the solution of the parameterized system is analyzed, leading to an additional optimality condition. By incorporating this optimality condition and the usual disconjugacy condition into the parameterized system, the dataset for optimal trajectories can be generated by propagating the parameterized system without using any optimization methods. In addition, a scaling invariance property is found for the solutions of the parameterized system. As a consequence of this scaling invariance property, a simple feedforward neural network trained by the solution of the parameterized system, selected at any fixed time, can be used to generate the nonlinear optimal guidance within milliseconds. Finally, numerical examples are presented, showing that the nonlinear optimal guidance command generated by the trained network can not only ensure the expected impact angle and impact time are precisely met but also requires less control effort compared with existing guidance methods. △ Less

Submitted 7 June, 2024; originally announced June 2024.

arXiv:2406.04690 [pdf, other]

Higher-order Structure Based Anomaly Detection on Attributed Networks

Authors: Xu Yuan, Na Zhou, Shuo Yu, Huafei Huang, Zhikui Chen, Feng Xia

Abstract: Anomaly detection (such as telecom fraud detection and medical image detection) has attracted the increasing attention of people. The complex interaction between multiple entities widely exists in the network, which can reflect specific human behavior patterns. Such patterns can be modeled by higher-order network structures, thus benefiting anomaly detection on attributed networks. However, due to… ▽ More Anomaly detection (such as telecom fraud detection and medical image detection) has attracted the increasing attention of people. The complex interaction between multiple entities widely exists in the network, which can reflect specific human behavior patterns. Such patterns can be modeled by higher-order network structures, thus benefiting anomaly detection on attributed networks. However, due to the lack of an effective mechanism in most existing graph learning methods, these complex interaction patterns fail to be applied in detecting anomalies, hindering the progress of anomaly detection to some extent. In order to address the aforementioned issue, we present a higher-order structure based anomaly detection (GUIDE) method. We exploit attribute autoencoder and structure autoencoder to reconstruct node attributes and higher-order structures, respectively. Moreover, we design a graph attention layer to evaluate the significance of neighbors to nodes through their higher-order structure differences. Finally, we leverage node attribute and higher-order structure reconstruction errors to find anomalies. Extensive experiments on five real-world datasets (i.e., ACM, Citation, Cora, DBLP, and Pubmed) are implemented to verify the effectiveness of GUIDE. Experimental results in terms of ROC-AUC, PR-AUC, and Recall@K show that GUIDE significantly outperforms the state-of-art methods. △ Less

Submitted 7 June, 2024; originally announced June 2024.

arXiv:2406.04594 [pdf, other]

Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach

Authors: Jianbo Dong, Bin Luo, Jun Zhang, Pengcheng Zhang, Fei Feng, Yikai Zhu, Ang Liu, Zian Chen, Yi Shi, Hairong Jiao, Gang Lu, Yu Guan, Ennan Zhai, Wencong Xiao, Hanyu Zhao, Man Yuan, Siran Yang, Xiang Li, Jiamang Wang, Rui Men, Jianwei Zhang, Huang Zhong, Dennis Cai, Yuan Xie, Binzhang Fu

Abstract: The emergence of Large Language Models (LLMs) has necessitated the adoption of parallel training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, we have found that the efficiency of current parallel training is often suboptimal, largely due to the following two main issues. Firstly, hardware failures are inevitable, leading to interruptions in the… ▽ More The emergence of Large Language Models (LLMs) has necessitated the adoption of parallel training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, we have found that the efficiency of current parallel training is often suboptimal, largely due to the following two main issues. Firstly, hardware failures are inevitable, leading to interruptions in the training tasks. The inability to quickly identify the faulty components results in a substantial waste of GPU resources. Secondly, since GPUs must wait for parameter synchronization to complete before proceeding to the next round of computation, network congestions can greatly increase the waiting time for GPUs. To address these challenges, this paper introduces a communication-driven solution, namely the C4. The key insights of C4 are two folds. First, in parallel training, collective communication exhibits periodic and homogeneous characteristics, so any anomalies are certainly due to some form of hardware malfunction. By leveraging this feature, C4 can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving few large flows, allows C4 to efficiently execute traffic planning, substantially reducing network congestion. C4 has been extensively implemented across our production systems, cutting error-induced overhead by roughly 30% and enhancing runtime performance by about 15% for certain applications with moderate communication costs. △ Less