CORTEXA
← Browse
crossrefFuture Internet2023-09-12Cited by 9

On Evaluating IoT Data Trust via Machine Learning

Timothy Tadj, Reza Arablouei, Volkan Dedeoglu

Data trust in IoT is crucial for safeguarding privacy, security, reliable decision-making, user acceptance, and complying with regulations. Various approaches based on supervised or unsupervised machine learning (ML) have recently been proposed for evaluating IoT data trust. However, assessing their real-world efficacy is hard mainly due to the lack of related publicly available datasets that can be used for benchmarking. Since obtaining such datasets is challenging, we propose a data synthesis method, called random walk infilling (RWI), to augment IoT time-series datasets by synthesizing untrustworthy data from existing trustworthy data. Thus, RWI enables us to create labeled datasets that can be used to develop and validate ML models for IoT data trust evaluation. We also extract new features from IoT time-series sensor data that effectively capture its autocorrelation as well as its cross-correlation with the data of the neighboring (peer) sensors. These features can be used to learn ML models for recognizing the trustworthiness of IoT sensor data. Equipped with our synthesized ground-truth-labeled datasets and informative correlation-based features, we conduct extensive experiments to critically examine various approaches to evaluating IoT data trust via ML. The results reveal that commonly used ML-based approaches to IoT data trust evaluation, which rely on unsupervised cluster analysis to assign trust labels to unlabeled data, perform poorly. This poor performance is due to the underlying assumption that clustering provides reliable labels for data trust, which is found to be untenable. The results also indicate that ML models, when trained on datasets augmented via RWI and using the proposed features, generalize well to unseen data and surpass existing related approaches. Moreover, we observe that a semi-supervised ML approach that requires only about 10% of the data labeled offers competitive performance while being practically more appealing compared to the fully supervised approaches. The related Python code and data are available online.

View free PDFSource page

Related papers

crossrefFuture Internet2023-06-09Cited by 11

Enhancing IoT Device Security through Network Attack Data Analysis Using Machine Learning Algorithms

Ashish Koirala, Rabindra Bista, Joao C. Ferreira

The Internet of Things (IoT) shares the idea of an autonomous system responsible for transforming physical computational devices into smart ones. Contrarily, storing and operating information and maintaining its confidentiality and security is a concerning issue in the IoT. Throu…

View free PDFSource page
crossrefFuture Internet2023-10-10Cited by 2

Data-Driven Safe Deliveries: The Synergy of IoT and Machine Learning in Shared Mobility

Fatema Elwy, Raafat Aburukba, A. R. Al-Ali, Ahmad Al Nabulsi, Alaa Tarek, Ameen Ayub, et al.

Shared mobility is one of the smart city applications in which traditional individually owned vehicles are transformed into shared and distributed ownership. Ensuring the safety of both drivers and riders is a fundamental requirement in shared mobility. This work aims to design a…

View free PDFSource page
crossrefFuture Internet2024-01-30Cited by 12

Context-Aware Behavioral Tips to Improve Sleep Quality via Machine Learning and Large Language Models

Erica Corda, Silvia M. Massa, Daniele Riboni

As several studies demonstrate, good sleep quality is essential for individuals’ well-being, as a lack of restoring sleep may disrupt different physical, mental, and social dimensions of health. For this reason, there is increasing interest in tools for the monitoring of sleep ba…

View free PDFSource page
crossrefFuture Internet2025-05-29

Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies

Qi Liu, Wanjing Ma

Data corruption, including missing and noisy entries, is a common challenge in real-world machine learning. This paper examines its impact and mitigation strategies through two experimental setups: supervised NLP tasks (NLP-SL) and deep reinforcement learning for traffic signal c…

View free PDFSource page
crossrefFuture Internet2024-05-12Cited by 23

Evaluating Realistic Adversarial Attacks against Machine Learning Models for Windows PE Malware Detection

Muhammad Imran, Annalisa Appice, Donato Malerba

During the last decade, the cybersecurity literature has conferred a high-level role to machine learning as a powerful security paradigm to recognise malicious software in modern anti-malware systems. However, a non-negligible limitation of machine learning methods used to train…

View free PDFSource page
crossrefFuture Internet2026-05-24

Enhancing the Adoption of Zero Trust in Organizations Using Machine Learning

Aeshah Mohammed Alshehri, Samer H. Atawneh, Hussein Al Bazar, Roxane Elias Mallouhy

Cybersecurity has become a critical concern for individuals, organizations, and governments, especially with the rise of sophisticated cyberattacks and remote work environments. Traditional security approaches are no longer sufficient, leading to the adoption of advanced framewor…

View free PDFSource page