Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Edward Sazonov

Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

Oct 02, 2024

Xinlong Hou, Sen Shen, Xueshen Li, Xinran Gao, Ziyi Huang, Steven J. Holiday, Matthew R. Cribbet, Susan W. White, Edward Sazonov, Yu Gan

Figure 1 for Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

Figure 2 for Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

Figure 3 for Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

Figure 4 for Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

Abstract:Being able to accurately monitor the screen exposure of young children is important for research on phenomena linked to screen use such as childhood obesity, physical activity, and social interaction. Most existing studies rely upon self-report or manual measures from bulky wearable sensors, thus lacking efficiency and accuracy in capturing quantitative screen exposure data. In this work, we developed a novel sensor informatics framework that utilizes egocentric images from a wearable sensor, termed the screen time tracker (STT), and a vision language model (VLM). In particular, we devised a multi-view VLM that takes multiple views from egocentric image sequences and interprets screen exposure dynamically. We validated our approach by using a dataset of children's free-living activities, demonstrating significant improvement over existing methods in plain vision language models and object detection models. Results supported the promise of this monitoring approach, which could optimize behavioral research on screen exposure in children's naturalistic settings.

* Prepare for submission

Via

Access Paper or Ask Questions

Improving Food Detection For Images From a Wearable Egocentric Camera

Jan 19, 2023

Yue Han, Sri Kalyan Yarlagadda, Tonmoy Ghosh, Fengqing Zhu, Edward Sazonov, Edward J. Delp

Figure 1 for Improving Food Detection For Images From a Wearable Egocentric Camera

Figure 2 for Improving Food Detection For Images From a Wearable Egocentric Camera

Figure 3 for Improving Food Detection For Images From a Wearable Egocentric Camera

Figure 4 for Improving Food Detection For Images From a Wearable Egocentric Camera

Abstract:Diet is an important aspect of our health. Good dietary habits can contribute to the prevention of many diseases and improve the overall quality of life. To better understand the relationship between diet and health, image-based dietary assessment systems have been developed to collect dietary information. We introduce the Automatic Ingestion Monitor (AIM), a device that can be attached to one's eye glasses. It provides an automated hands-free approach to capture eating scene images. While AIM has several advantages, images captured by the AIM are sometimes blurry. Blurry images can significantly degrade the performance of food image analysis such as food detection. In this paper, we propose an approach to pre-process images collected by the AIM imaging sensor by rejecting extremely blurry images to improve the performance of food detection.

* 6 pages, 6 figures, Conference Paper for Imaging and Multimedia Analytics in a Web and Mobile World Conference, IS&T Electronic Imaging Symposium, Burlingame, CA (Virtual), January, 2021

Via

Access Paper or Ask Questions

Clustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning

Aug 25, 2022

Jiachuan Peng, Peilun Shi, Jianing Qiu, Xinwei Ju, Frank P. -W. Lo, Xiao Gu, Wenyan Jia, Tom Baranowski, Matilda Steiner-Asiedu, Alex K. Anderson(+5 more)

Figure 1 for Clustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning

Figure 2 for Clustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning

Figure 3 for Clustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning

Figure 4 for Clustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning

Abstract:In our recent dietary assessment field studies on passive dietary monitoring in Ghana, we have collected over 250k in-the-wild images. The dataset is an ongoing effort to facilitate accurate measurement of individual food and nutrient intake in low and middle income countries with passive monitoring camera technologies. The current dataset involves 20 households (74 subjects) from both the rural and urban regions of Ghana, and two different types of wearable cameras were used in the studies. Once initiated, wearable cameras continuously capture subjects' activities, which yield massive amounts of data to be cleaned and annotated before analysis is conducted. To ease the data post-processing and annotation tasks, we propose a novel self-supervised learning framework to cluster the large volume of egocentric images into separate events. Each event consists of a sequence of temporally continuous and contextually similar images. By clustering images into separate events, annotators and dietitians can examine and analyze the data more efficiently and facilitate the subsequent dietary assessment processes. Validated on a held-out test set with ground truth labels, the proposed framework outperforms baselines in terms of clustering quality and classification accuracy.

* accepted to BHI 2022

Via

Access Paper or Ask Questions

Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

Jul 01, 2021

Jianing Qiu, Frank P. -W. Lo, Xiao Gu, Modou L. Jobarteh, Wenyan Jia, Tom Baranowski, Matilda Steiner-Asiedu, Alex K. Anderson, Megan A McCrory, Edward Sazonov(+3 more)

Figure 1 for Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

Figure 2 for Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

Figure 3 for Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

Figure 4 for Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

Abstract:Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eating behaviours of the subject. However, there currently is no method that is able to incorporate these visual clues and provide a comprehensive context of dietary intake from passive recording (e.g., is the subject sharing food with others, what food the subject is eating, and how much food is left in the bowl). On the other hand, privacy is a major concern while egocentric wearable cameras are used for capturing. In this paper, we propose a privacy-preserved secure solution (i.e., egocentric image captioning) for dietary assessment with passive monitoring, which unifies food recognition, volume estimation, and scene understanding. By converting images into rich text descriptions, nutritionists can assess individual dietary intake based on the captions instead of the original images, reducing the risk of privacy leakage from images. To this end, an egocentric dietary image captioning dataset has been built, which consists of in-the-wild images captured by head-worn and chest-worn cameras in field studies in Ghana. A novel transformer-based architecture is designed to caption egocentric dietary images. Comprehensive experiments have been conducted to evaluate the effectiveness and to justify the design of the proposed architecture for egocentric dietary image captioning. To the best of our knowledge, this is the first work that applies image captioning to dietary intake assessment in real life settings.

Via

Access Paper or Ask Questions