Explore the latest peer-reviewed research with AI-generated summaries, key findings, and insights for easy understanding. All content is legally sourced from academic metadata with links to original papers.
Tracking often uses a multistage pipeline of feature extraction, target information integration, and bounding box estimation. We also perform in-depth ablation studies to demonstrate the effectiveness...
Human skeleton, as a compact representation of human action, has received increasing attention in recent years. Once fused with other modalities, it achieves the advanced on all eight multi-modality a...
Vision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Our model is currently the a...
Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Experiments on five challengin...
Knowledge distillation (KD) achieves promising results on the challenging problem of unsupervised anomaly detection (AD). The obtained compact embedding effectively preserves essential information on...
In this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection. Without bells-and-whistles, MViTv2 has advanc...
We present Block-NeRF, a variant of Neural Radiance Fields that can represent large-scale environments. We add appearance embeddings, learned pose refinement, and controllable exposure to each individ...
With the rapid increase in the volume of data on the aquatic environment, machine learning has become an important tool for data analysis, classification, and prediction. In this review, we describe t...
Retinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. Extensive experiments on real-world low-light images qualit...
We present Point-BERT, a new paradigm for learning Transformers to generalize the concept of BERT to 3D point cloud. We also demonstrate that the representations learned by Point-BERT transfer well to...
Natural language offers a highly intuitive interface for image editing. We compare against several baselines and related methods, both qualitatively and quantitatively, and show that our method outper...
This research explores on-chip photonic deep neural network for image classificatio..., contributing new insights to the field of Artificial Intelligence.
The mainstream paradigm behind continual learning has been to adapt the model parameters to non-stationary data distributions, where catastrophic forgetting is the central challenge. Surprisingly, L2P...
For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other. Moreover, we estab...
We present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. Our experiments indicate that the VQ-Diffusion model with the reparameterization is fifteen times faster th...
Single image super-resolution (SISR) has witnessed great strides with the development of deep learning. Compared with the original Transformer which occupies 16,057M GPU memory, ESRT only occupies 4,1...
We present Mobile-Former, a parallel design of MobileNet and transformer with a two-way bridge in between. Additionally, we build an efficient end-to-end detector by replacing backbone, encoder and de...
This research explores Optical vegetation indices for monitoring terrestrial ecosys..., contributing new insights to the field of Artificial Intelligence.
Machine learning (ML) is a new-age thriving technology, which facilitates computers to read and interpret from the previously present data automatically. The paper answers all relevant questions that...
Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars. The ver...