Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Ege Ozguroglu

EraseDraw: Learning to Insert Objects by Erasing Them from Images

Aug 31, 2024

Alper Canberk, Maksym Bondarenko, Ege Ozguroglu, Ruoshi Liu, Carl Vondrick

Figure 1 for EraseDraw: Learning to Insert Objects by Erasing Them from Images

Figure 2 for EraseDraw: Learning to Insert Objects by Erasing Them from Images

Figure 3 for EraseDraw: Learning to Insert Objects by Erasing Them from Images

Figure 4 for EraseDraw: Learning to Insert Objects by Erasing Them from Images

Abstract:Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often fail by making global changes to the image, inserting objects in unrealistic spatial locations, and generating inaccurate lighting details. We observe that while state-of-the-art models perform poorly on object insertion, they can remove objects and erase the background in natural images very well. Inverting the direction of object removal, we obtain high-quality data for learning to insert objects that are spatially, physically, and optically consistent with the surroundings. With this scalable automatic data generation pipeline, we can create a dataset for learning object insertion, which is used to train our proposed text conditioned diffusion model. Qualitative and quantitative experiments have shown that our model achieves state-of-the-art results in object insertion, particularly for in-the-wild images. We show compelling results on diverse insertion prompts and images across various domains.In addition, we automate iterative insertion by combining our insertion model with beam search guided by CLIP.

Via

Access Paper or Ask Questions

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Jun 24, 2024

Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, Carl Vondrick

Figure 1 for Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Figure 2 for Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Figure 3 for Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Figure 4 for Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Abstract:A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a visuomotor policy learning framework that fine-tunes a video diffusion model on human demonstrations of a given task. At test time, we generate an example of an execution of the task conditioned on images of a novel scene, and use this synthesized execution directly to control the robot. Our key insight is that using common tools allows us to effortlessly bridge the embodiment gap between the human hand and the robot manipulator. We evaluate our approach on four tasks of increasing complexity and demonstrate that harnessing internet-scale generative models allows the learned policy to achieve a significantly higher degree of generalization than existing behavior cloning approaches.

* Project page: https://dreamitate.cs.columbia.edu/

Via

Access Paper or Ask Questions

Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

May 23, 2024

Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sargent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, Carl Vondrick

Figure 1 for Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

Figure 2 for Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

Figure 3 for Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

Figure 4 for Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis

Abstract:Accurate reconstruction of complex dynamic scenes from just a single viewpoint continues to be a challenging task in computer vision. Current dynamic novel view synthesis methods typically require videos from many different camera viewpoints, necessitating careful recording setups, and significantly restricting their utility in the wild as well as in terms of embodied AI applications. In this paper, we propose $\textbf{GCD}$, a controllable monocular dynamic view synthesis pipeline that leverages large-scale diffusion priors to, given a video of any scene, generate a synchronous video from any other chosen perspective, conditioned on a set of relative camera pose parameters. Our model does not require depth as input, and does not explicitly model 3D scene geometry, instead performing end-to-end video-to-video translation in order to achieve its goal efficiently. Despite being trained on synthetic multi-view video data only, zero-shot real-world generalization experiments show promising results in multiple domains, including robotics, object permanence, and driving environments. We believe our framework can potentially unlock powerful applications in rich dynamic scene understanding, perception for robotics, and interactive 3D video viewing experiences for virtual reality.

* Project webpage is available at: https://gcd.cs.columbia.edu/

Via

Access Paper or Ask Questions

pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Jan 25, 2024

Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, Carl Vondrick

Figure 1 for pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Figure 2 for pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Figure 3 for pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Figure 4 for pix2gestalt: Amodal Segmentation by Synthesizing Wholes

Abstract:We introduce pix2gestalt, a framework for zero-shot amodal segmentation, which learns to estimate the shape and appearance of whole objects that are only partially visible behind occlusions. By capitalizing on large-scale diffusion models and transferring their representations to this task, we learn a conditional diffusion model for reconstructing whole objects in challenging zero-shot cases, including examples that break natural and physical priors, such as art. As training data, we use a synthetically curated dataset containing occluded objects paired with their whole counterparts. Experiments show that our approach outperforms supervised baselines on established benchmarks. Our model can furthermore be used to significantly improve the performance of existing object recognition and 3D reconstruction methods in the presence of occlusions.

* Website: https://gestalt.cs.columbia.edu/

Via

Access Paper or Ask Questions