Explore the latest peer-reviewed research with AI-generated summaries, key findings, and insights for easy understanding. All content is legally sourced from academic metadata with links to original papers.
Free-form inpainting is the task of adding new content to an image in the regions specified by an arbitrary binary mask. Re-Paint outperforms advanced Autoregressive, and GAN approaches for at least f...
Being able to spot defective parts is a critical component in large-scale industrial manufacturing. We further report competitive results on two additional datasets and also find competitive results i...
We revisit large kernel design in modern convolutional neural networks (CNNs). Our study further reveals that, in contrast to small-kernel CNNs, large-kernel CNNs have much larger effective receptive...
The ability to learn richer network representations generally boosts the performance of deep learning models. Adding a Split-Attention module into the architecture design space of RegNet-Y and FBNetV2...
We introduce Plenoxels (plenoptic voxels), a systemfor photorealistic view synthesis. On standard, benchmark tasks, Plenoxels are optimized two orders of magnitude faster than Neural Radiance Fields w...
We present CSWin Transformer, an efficient and effective Transformer-based backbone for general-purpose vision tasks. By further pretraining on the larger dataset ImageNet-21K, we achieve 87.5% Top-1...
This paper presents SimMIM, a simple framework for masked image modeling. We also leverage this approach to address the data-hungry issue faced by large-scale model training, that a 3B model (Swin V2-...
Transformers have shown great potential in computer vision tasks. This work calls for more future research dedicated to improving MetaFormer instead of focusing on the token mixer modules.
Unsupervised generation of high-quality multi-view-consistent images and 3D shapes using only collections of single-view 2D photographs has been a long-standing challenge. By decoupling feature genera...
This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Extensive experiments on several public datasets and our benchmark demonstrate that o...
Existing low-light image enhancement techniques are mostly not only difficult to deal with both visual quality and computational efficiency but also commonly invalid in unknown complex scenarios. Appl...
The challenging task of multi-object tracking (MOT) requires simultaneous reasoning about track initialization, identity, and spatio-temporal trajectories. TrackFormer introduces a new tracking-by-att...
We present in this paper a novel denoising training method to speedup DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. Compa...
We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. Finally, evaluation on five inward-facing benc...
advanced distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. This paper proves the great pot...
Transformers have recently shown superior performances on various vision tasks. Extensive experi-ments show that our models achieve consistently improved results on comprehensive benchmarks.
Vision transformers have been successfully applied to image recognition tasks due to their ability to capture long-range dependencies within an image. In particular, our CMT-S achieves 83.5% top-1 acc...
LiDAR and camera are two important sensors for 3D object detection in autonomous driving. We provide extensive experiments to demonstrate its robustness against degenerated image quality and calibrati...
Attention-based neural networks such as the Vision Transformer (ViT) have recently attained advanced results on many computer vision benchmarks. As a result, we successfully train a ViT model with two...
A commonly observed failure mode of Neural Radiance Field (NeRF) is fitting incorrect geometries when given an insufficient number of input views. Further, we show that our loss is compatible with oth...