Privacy Risk Predictions Based on Fundamental Understanding of Personal Data and an Evolving Threat Landscape
Haoran Niu, K. Suzanne Barber, UTCID Report #26-03, March 2026
Privacy Risk Predictions Based on Fundamental Understanding of Personal Data and an Evolving Threat Landscape
Haoran Niu, K. Suzanne Barber, UTCID Report #26-03, March 2026
Show Abstract
It is difficult for individuals and organizations to protect personal information without a fundamental understanding of relative privacy risks. By analyzing over 5,000 empirical identity theft and fraud cases, this research identifies which types of personal data are exposed, how frequently such exposures occur, and what the consequences of those exposures are. We construct an Identity Ecosystem graph—a foundational, graph-based model in which nodes represent personally identifiable information (PII) attributes and edges represent empirical disclosure relationships between them (e.g., one PII attribute is exposed due to the exposure of another). Leveraging this graph structure, we develop a privacy risk prediction framework that uses graph theory and graph neural networks to estimate the likelihood of further disclosures when certain PII attributes are compromised. The results show that our approach effectively addresses the core question: Can the disclosure of a given identity attribute possibly lead to the disclosure of another attribute?
Access Publication: Download PDF of Report
A Survey on Quantitative Modeling of Trust in Online Social Networks
Wenting Song, K. Suzanne Barber, UTCID Report #26-01, February 2026
A Survey on Quantitative Modeling of Trust in Online Social Networks
Wenting Song, K. Suzanne Barber, UTCID Report #26-01, February 2026
Show Abstract
Online social networks facilitate user engagement and information sharing but are also rife with misinformation and deception. Research on trust modeling in online social networks focuses on developing computational models or algorithms to measure trust relationships, assess the reliability of shared content, and detect spam or malicious activities. However, most existing review papers either briefly mention the concept of trust or focus on a single category of trust models. In this paper, we offer a comprehensive categorization and review of state-of-the-art trust models developed for online social networks. First, we explore theories and modelsrelated to trust in psychology and identify several factors that influence the formation and evolution of online trust. Next, state-of-the-art trust models are categorized based on their algorithmic foundations. For each category, the modeling mechanisms are investigated, and their unique contributions to quantitative trust modeling are highlighted. Subsequently, we provide an implementation-centric trust modeling handbook, which summarizes available datasets, trust-related features, promising modeling techniques, and feasible application scenarios. Finally, the findings of the literature review are summarized, and unresolved challenges are discussed.
Access Publication: Download PDF of Report
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
Jeffri Murrugarra-Llerena, Haoran Niu, K. Suzanne Barber, Hal Daume III, Trista Cao, Paola Cascante-Bonilla, UT CID Report #25-05, October 2025
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
Jeffri Murrugarra-Llerena, Haoran Niu, K. Suzanne Barber, Hal Daume III, Trista Cao, Paola Cascante-Bonilla, UT CID Report #25-05, October 2025
Show Abstract
As visual assistant systems powered by visual language models (VLMs) become more prevalent, concerns over user privacy have grown, particularly for blind and low vision users who may unknowingly capture personal private information in their images. Existing privacy protection methods rely on coarse-grained segmentation, which uniformly masks entire private objects, often at the cost of usability. In this work, we propose FiG-Priv, a fine-grained privacy protection framework that selectively masks only high-risk private information while preserving low-risk information. Our approach integrates fine-grained segmentation with a data-driven risk scoring mechanism. By leveraging a more nuanced understanding of privacy risk, our method enables more effective protection without unnecessarily restricting users’ access to critical information. We evaluate our framework using the BIV-Priv-Seg dataset and show that FiG-Priv preserves +26% of image content, enhancing the ability of VLMs to provide useful responses by 11% and identify the image content by 45%, while ensuring privacy protection.
Access Publication: Download PDF of Report
A Longitudinal Look at GDPR Compliance
Brian Kim, Trista Cao, K.Suzanne Barber, UT CID Report #25-08, August 2025
A Longitudinal Look at GDPR Compliance
Brian Kim, Trista Cao, K.Suzanne Barber, UT CID Report #25-08, August 2025
Show Abstract
This paper presents a longitudinal study investigating how the General Data Protection Regulation (GDPR) compliance of website privacy policies has evolved over a five-yearperiod. Using an automated privacy policy evaluation tool, we assessed ten core GDPR factors across a corpus of websites originally analyzed in 2020 and re-evaluated in 2025. Our analysis reveals a mixed progression: while user-facing compliance measures such as consent, data retention notification, and data sharing transparency showed measurable improvement, technically oriented factors—such as breach notification and data encryption—experienced a decline in explicit disclosure. These findings suggest a broader trend in which privacy policies increasingly emphasize legal rights and visible consent mechanisms, while de-emphasizing backend technical safeguards. The results point to a split in compliance communication, possibly influenced by regulatory clarity, enforcement pressure, and shifts in organizational privacy strategy. This study underscores the importance of continued policy auditing and the need for complementary methods that bridge the gap between stated policy and implemented practice in the context of evolving digital governance frameworks.
Access Publication: Download PDF of Report
TrustLabTM: An Interactive Tool for Evaluating Online Trustworthiness Across Diverse Domains
Teng-Chieh Huang, K.Suzanne Barber, UT CID Report #25-03, March 2025
TrustLabTM: An Interactive Tool for Evaluating Online Trustworthiness Across Diverse Domains
Teng-Chieh Huang, K.Suzanne Barber, UT CID Report #25-03, March 2025
Show Abstract
TrustLabTM is an innovative online tool for assessing social media users’ trustworthiness with ease and precision. It serves diverse users, from general social media participants to researchers aiming to gauge trust levels in various domains. Unlike many tools, TrustLabTM focuses on user trustworthiness rather than post content, distinguishing between experts and typical users. Using Trust Filters and user attributes, it assigns trust scores visualized through intuitive charts for clarity. Additionally, TrustLabTM provides personalized recommendations to help users enhance their online credibility. While its algorithms are domain-independent, this paper demonstrates TrustLabTM’s application in finance, politics, and health, showcasing its role in shaping public discourse, knowledge, and connections. With its user-friendly interface, TrustLabTM is a significant tool for exploring and understanding online trust in the digital era.
Access Publication: Download PDF of Report
TWCF: Trust Weighted Collaborative Filtering based on Quantitative Modeling of Trust
Wenting Song, K. Suzanne Barber, UT CID Report #24-10, December 2024
TWCF: Trust Weighted Collaborative Filtering based on Quantitative Modeling of Trust
Wenting Song, K. Suzanne Barber, UT CID Report #24-10, December 2024
Show Abstract
Trust-based collaborative filtering methods exploit trust information to alleviate cold-start and data sparsity prob-lems and improve recommendation performance. Instead of modeling trust itself, most existing approaches rely on trust connections sourced from online social networking platforms. Typically, users engage with products on rating platforms and interact with friends on social networking platforms, exhibit-ing distinct behavior patterns. Therefore, simply aggregating information and misinterpreting user preferences may lead to information misuse and recommendation bias. To overcome these limitations, this paper introduces three trust measurement models to capture and quantify trust relationships directly from user-provided ratings. Utilizing the trust-weighted scheme, we propose a hybrid collaborative filtering approach called Trust Weighted Collaborative Filtering (TWCF). Experiments on three real-world datasets show that the trust information extracted from rating data and the trust-weighted scheme can significantly improve the performance of original neighborhood collaborative filtering. TWCF achieves an average improvement of 7.12% in prediction accuracy over the trust-based collaborative filtering baselines. Furthermore, the interpretability and scalability of TWCF provide opportunities for further improvement by incorporating more appropriate trust measurement models.
Access Publication: Download PDF of Report
Using Historical Social Media Retrieved Trust Attributes to Help Distinguishing Trustworthy Users
Teng-Chieh Huang, Razieh Nokhbeh Zaemm, K.Suzanne Barber, UT CID Report #23-04, March 2023
Using Historical Social Media Retrieved Trust Attributes to Help Distinguishing Trustworthy Users
Teng-Chieh Huang, Razieh Nokhbeh Zaemm, K.Suzanne Barber, UT CID Report #23-04, March 2023
Show Abstract
With the penetration of social media across the world, the information generated by the users has increased exponentially. The wisdom of crowds can now be easily accessed from the Internet. The problem is, how to correctly interpret the true public opinion without distorting it? Considering the spam or malicious users hidden in the social media spreading misinformation and disinformation, the solution might not be trivial. Some previous work uses quantity accumulation trying to mitigate the influence of bad users. More recent works add machine learning techniques to help with the correct judgment. However, the importance of the individual user - the actual person who is behind the screen, does not attract the attention it deserves. In this work, focusing on the history of user behavior, we provide a different angle to understand the connection between the credibility of social media users and the trustworthiness of their virtual representatives. By analyzing Twitter data from November 2017 to November 2021 which contains three types of users (typical, topic-related, and expert) on two target domains (politics and finance), we can gain deeper insights on the social media users and their trustworthiness.
Access Publication: Download PDF of Report
Social Media Trustworthy User Identification Across Politics and Finance Domains
Razieh Nokhbeh Zaeem, K. Suzanne Barber, Kai Chih Chang, UT CID Report #22-07, July 2022
Social Media Trustworthy User Identification Across Politics and Finance Domains
Razieh Nokhbeh Zaeem, K. Suzanne Barber, Kai Chih Chang, UT CID Report #22-07, July 2022
Show Abstract
This research seeks to answer two ever-pressing questions for those who rely on social media for information – What information can I trust? and Who can I trust? Specifically, this research addresses how to recommend trustworthy information from specific information sources using an innovative method of combining interpersonal trust on the Internet and classification algorithms for various target domains. A variety of domains have been studied to see how information from social media can predict or explain the phenomena of the domain, from politics (who wins the next election) to finance (what happens to a company’s stock). It is important, however, that as we improve our methods to interpret the information retrieved from social media, we also pay attention to the quality of that information. Social media information is noisy and possibly contains useless, misleading, or even malicious information distributed by untrustworthy users. It is paramount to filter credible and trustworthy information generated by trustworthy users and domain experts from contaminated data, advertisements, or scams. In this paper, we develop a novel, domain-independent method to distinguish three types of social media (specifically Twitter) users from one another: typical users, domain-related users, and experts. We take a comprehensive list of trust attributes (i.e., measurable properties of a social media user account or posts such as the number of tweets or the number of followers) from previous work and study the value of these trust attributes for the three types of users. The value of the trust attribute for typical users is the average value for all the users. Domain-related users are those who tweet with a set of domain-related handles (e.g., @a_stock_ticker). Experts are real world experts of the filed, who we extract from reputable journalistic sources outside social media. By applying random forest to compare trust attribute performance, we identify which trust attributes can best distinguish expert, domain-related, and typical users from one another. Most importantly, we keep our work independent of the subject domain by performing the same set of experiments in the finance as well as the politics domain. We compare the distributions of trust attributes between finance and politics domains, to test if the classification method can be applied to various target domains and still provide robust results. Overall, we identify trust attributes capable of distinguishing trustworthy users from malicious ones, or to act as trust filters. Our work increases the reliability and utility of social media data for better decision-making. By applying trust filters to filter out untrustworthy and malicious users, the bad impact of social media such as fake news or media manipulation can thus be minimized.
Access Publication: Download PDF of Report
An Identity Asset Sensitivity Model in Self-Sovereign Identities
Razieh Nokhbeh Zaeem, K. Suzanne Barber, Kai Chih Chang, UT CID Report #21-07, August 2021
An Identity Asset Sensitivity Model in Self-Sovereign Identities
Razieh Nokhbeh Zaeem, K. Suzanne Barber, Kai Chih Chang, UT CID Report #21-07, August 2021
Show Abstract
Due to the emergence of new paradigms such as social media and the Internet of Things (IoT), the use of the Internet has ushered in further challenges. After years of research, there is still no com-plete layer of identity on the Internet. In order to provide identity management, self-sovereign identity has become a popular choice. Self-sovereign identities provide users with complete autonomy and immutability for personal identities, as well as complete control for their identity owners. Like any type of identity, a self-sovereign identity also processes the Personally Identifiable Information (PII) of the identity holder and faces privacy and security risks com-mon to identity management. This research proposes a model of determining PII sensitivity by a score to measure what attributes or combination thereof is sensitive to share. Our work highlights that while it is important to improve how PII attributes are shared, it is paramount to identify which PII attributes are safer to share to achieve the same identity management goals.
Access Publication: Download PDF of Report
Identifying Real-World Credible Experts in the Financial Domain
Teng-Chieh Huang, Razieh Nokhbeh Zaemm, Suzanne Barber, UT CID Report #21-11, December 2021
Identifying Real-World Credible Experts in the Financial Domain
Teng-Chieh Huang, Razieh Nokhbeh Zaemm, Suzanne Barber, UT CID Report #21-11, December 2021
Show Abstract
Establishing a solid mechanism for finding credible and trustworthy people in online social networks is an important first step to avoid useless, misleading or even malicious information. There is a body of existing work studying trustworthiness of social media users and finding credible sources in specific target domains. However, most of the related work lack the connection between the credibility in the real-world and credibility on the Internet, which makes the formation of social media credibility and trustworthiness incomplete. In this paper, working in the financial domain, we identify attributes that can distinguish credible users on the Internet who are indeed trustworthy experts in the real-world. To ensure objectivity, we gather the list of credible financial experts from real-world financial authorities. We analyze the distribution of attributes of about 10K stock-related Twitter users and their 600K tweets over six months in 2015/2016, and over 2.6M typical Twitter users and their 4.8M tweets on November 2nd, 2015, comprising 1% of the entire Twitter in that time period. By using the random forest classifier, we find which attributes are related to real-world expertise. Our work sheds light on the properties of trustworthy users and paves the way for their automatic identification.
Access Publication: Download PDF of Report