Oriented-grid Encoder for 3D Implicit Representations
Encoding 3D points is one of the primary steps in learning-based implicit scene representation. Using features that gather information from neighbors with multi-resolution grids has proven to be the...
View ArticleTransfer learning with generative models for object detection on limited...
The availability of data is limited in some fields, especially for object detection tasks, where it is necessary to have correctly labeled bounding boxes around each object. A notable example of such...
View ArticleIs it safe to cross? Interpretable Risk Assessment with GPT-4V for...
Safely navigating street intersections is a complex challenge for blind and low-vision individuals, as it requires a nuanced understanding of the surrounding context - a task heavily reliant on visual...
View ArticleFingerprinting New York City's Scaffolding Problem with Longitudinal Dashcam...
Scaffolds, also called sidewalk sheds, are intended to be temporary structures to protect pedestrians from construction and repair hazards. However, some sidewalk sheds are left up for years. Long-term...
View ArticleEvent-to-Video Conversion for Overhead Object Detection
Collecting overhead imagery using an event camera is desirable due to the energy efficiency of the image sensor compared to standard cameras. However, event cameras complicate downstream image...
View ArticleDomain Adaptation Using Pseudo Labels
In the absence of labeled target data, unsupervised domain adaptation approaches seek to align the marginal distributions of the source and target domains in order to train a classifier for the target....
View ArticleNeural Rendering based Urban Scene Reconstruction for Autonomous Driving
Dense 3D reconstruction has many applications in automated driving including automated annotation validation, multimodal data augmentation, providing ground truth annotations for systems lacking LiDAR,...
View ArticleGyroscope-Assisted Motion Deblurring Network
Image research has shown substantial attention in deblurring networks in recent years. Yet, their practical usage in real-world deblurring, especially motion blur, remains limited due to the lack of...
View ArticleLatent Enhancing AutoEncoder for Occluded Image Classification
Large occlusions result in a significant decline in image classification accuracy. During inference, diverse types of unseen occlusions introduce out-of-distribution data to the classification model,...
View ArticleSemantic Object-level Modeling for Robust Visual Camera Relocalization
Visual relocalization is crucial for autonomous visual localization and navigation of mobile robotics. Due to the improvement of CNN-based object detection algorithm, the robustness of visual...
View ArticleSynthesizing CTA Image Data for Type-B Aortic Dissection using Stable...
Stable Diffusion (SD) has gained a lot of attention in recent years in the field of Generative AI thus helping in synthesizing medical imaging data with distinct features. The aim is to contribute to...
View ArticleTreatment-wise Glioblastoma Survival Inference with Multi-parametric...
In this work, we aim to predict the survival time (ST) of glioblastoma (GBM) patients undergoing different treatments based on preoperative magnetic resonance (MR) scans. The personalized and precise...
View ArticleOSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted Surgery
In the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical activity recognition...
View ArticleReciprocal Visibility
We propose a guidance strategy to optimize real-time synthetic aperture sampling for occlusion removal with drones by pre-scanned point-cloud data. Depth information can be used to compute visibility...
View ArticleA Change Detection Reality Check
In recent years, there has been an explosion of proposed change detection deep learning architectures in the remote sensing literature. These approaches claim to offer state-of the-art performance on...
View ArticleDomain Adaptable Fine-Tune Distillation Framework For Advancing Farm...
In this study, we propose an automated framework for camel farm monitoring, introducing two key contributions: the Unified Auto-Annotation framework and the Fine-Tune Distillation framework. The...
View ArticleA Highlight Removal Method for Capsule Endoscopy Images
The images captured by Wireless Capsule Endoscopy (WCE) always exhibit specular reflections, and removing highlights while preserving the color and texture in the region remains a challenge. To address...
View ArticleA Benchmark for Multi-modal Foundation Models on Low-level Vision: from...
The rapid development of Multi-modality Large Language Models (MLLMs) has navigated a paradigm shift in computer vision, moving towards versatile foundational models. However, evaluating MLLMs in...
View ArticleTwo-Stage Multi-task Self-Supervised Learning for Medical Image Segmentation
Medical image segmentation has been significantly advanced by deep learning (DL) techniques, though the data scarcity inherent in medical applications poses a great challenge to DL-based segmentation...
View ArticleINSITE: labelling medical images using submodular functions and...
The necessity of large amounts of labeled data to train deep models, especially in medical imaging creates an implementation bottleneck in resource-constrained settings. In Insite (labelINg medical...
View Article