Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Guruprasad V Ramesh

Synthetic Counterfactual Faces

Jul 18, 2024

Guruprasad V Ramesh, Harrison Rosenberg, Ashish Hooda, Kassem Fawaz

Figure 1 for Synthetic Counterfactual Faces

Figure 2 for Synthetic Counterfactual Faces

Figure 3 for Synthetic Counterfactual Faces

Figure 4 for Synthetic Counterfactual Faces

Abstract:Computer vision systems have been deployed in various applications involving biometrics like human faces. These systems can identify social media users, search for missing persons, and verify identity of individuals. While computer vision models are often evaluated for accuracy on available benchmarks, more annotated data is necessary to learn about their robustness and fairness against semantic distributional shifts in input data, especially in face data. Among annotated data, counterfactual examples grant strong explainability characteristics. Because collecting natural face data is prohibitively expensive, we put forth a generative AI-based framework to construct targeted, counterfactual, high-quality synthetic face data. Our synthetic data pipeline has many use cases, including face recognition systems sensitivity evaluations and image understanding system probes. The pipeline is validated with multiple user studies. We showcase the efficacy of our face generation pipeline on a leading commercial vision model. We identify facial attributes that cause vision systems to fail.

* Paper under review. Full text and results will be updated after acceptance

Via

Access Paper or Ask Questions

Unbiased Face Synthesis With Diffusion Models: Are We There Yet?

Sep 13, 2023

Harrison Rosenberg, Shimaa Ahmed, Guruprasad V Ramesh, Ramya Korlakai Vinayak, Kassem Fawaz

Figure 1 for Unbiased Face Synthesis With Diffusion Models: Are We There Yet?

Figure 2 for Unbiased Face Synthesis With Diffusion Models: Are We There Yet?

Figure 3 for Unbiased Face Synthesis With Diffusion Models: Are We There Yet?

Figure 4 for Unbiased Face Synthesis With Diffusion Models: Are We There Yet?

Abstract:Text-to-image diffusion models have achieved widespread popularity due to their unprecedented image generation capability. In particular, their ability to synthesize and modify human faces has spurred research into using generated face images in both training data augmentation and model performance assessments. In this paper, we study the efficacy and shortcomings of generative models in the context of face generation. Utilizing a combination of qualitative and quantitative measures, including embedding-based metrics and user studies, we present a framework to audit the characteristics of generated faces conditioned on a set of social attributes. We applied our framework on faces generated through state-of-the-art text-to-image diffusion models. We identify several limitations of face image generation that include faithfulness to the text prompt, demographic disparities, and distributional shifts. Furthermore, we present an analytical model that provides insights into how training data selection contributes to the performance of generative models.

Via

Access Paper or Ask Questions

Federated Representation Learning for Automatic Speech Recognition

Aug 07, 2023

Guruprasad V Ramesh, Gopinath Chennupati, Milind Rao, Anit Kumar Sahu, Ariya Rastrow, Jasha Droppo

Figure 1 for Federated Representation Learning for Automatic Speech Recognition

Figure 2 for Federated Representation Learning for Automatic Speech Recognition

Figure 3 for Federated Representation Learning for Automatic Speech Recognition

Figure 4 for Federated Representation Learning for Automatic Speech Recognition

Abstract:Federated Learning (FL) is a privacy-preserving paradigm, allowing edge devices to learn collaboratively without sharing data. Edge devices like Alexa and Siri are prospective sources of unlabeled audio data that can be tapped to learn robust audio representations. In this work, we bring Self-supervised Learning (SSL) and FL together to learn representations for Automatic Speech Recognition respecting data privacy constraints. We use the speaker and chapter information in the unlabeled speech dataset, Libri-Light, to simulate non-IID speaker-siloed data distributions and pre-train an LSTM encoder with the Contrastive Predictive Coding framework with FedSGD. We show that the pre-trained ASR encoder in FL performs as well as a centrally pre-trained model and produces an improvement of 12-15% (WER) compared to no pre-training. We further adapt the federated pre-trained models to a new language, French, and show a 20% (WER) improvement over no pre-training.

* Accepted at ISCA SPSC Symposium 3rd Symposium on Security and Privacy in Speech Communication, 2023

Via

Access Paper or Ask Questions