Explore the latest peer-reviewed research with AI-generated summaries, key findings, and insights for easy understanding. All content is legally sourced from academic metadata with links to original papers.
While the vast majority of well-structured single protein chains can now be predicted to high accuracy due to the recent AlphaFold model, the prediction of multi-chain protein complexes remains a chal...
This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. The hierarchical design and the shifted window approach al...
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). We implement our findings into...
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction task...
Image restoration is a long-standing low-level vision problem that aims to restore high-quality images from low-quality images (e.g., downscaled, noisy and compressed images). We conduct experiments o...
Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. For example, T2T-ViT with comp...
Self-attention networks have revolutionized natural language processing and are making impressive strides in image analysis tasks such as image classification and object detection. Our Point Transform...
Object detection on drone-captured scenarios is a recent popular task. On VisDrone Challenge 2021, TPH-YOLOv5 wins 5 <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/...
The recently developed vision transformer (ViT) has achieved promising results on image classification compared to convolutional neural networks. For example, on the ImageNet1K dataset, with some arch...
Though many attempts have been made in blind super-resolution to restore low-resolution images with unknown and complex degradations, they are still far from addressing general real-world degraded ima...
This paper does not describe a novel method. We discuss the currently positive evidence as well as challenges and open questions.
We introduce four new real-world distribution shift datasets consisting of changes in image style, image blurriness, geographic location, camera operation, and more. Overall we find that some methods...
Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. Combining CTR-GC with temporal modeling modules, we develop a powerful g...
This research explores physics of higher-order interactions in complex systems, contributing new insights to the field of Artificial Intelligence.
Visual surface anomaly detection aims to detect local image regions that significantly deviate from normal appearance. On the challenging MVTec anomaly detection dataset, DRÆM outperforms the current...
We introduce a method to render Neural Radiance Fields (NeRFs) in real time using PlenOctrees, an octree-based 3D representation which supports view-dependent effects. Our real-time neural rendering a...
Artificial intelligence (AI) has an astonishing potential in assisting clinical decision making and revolutionizing the field of health care. If the training data is misrepresentative of the populatio...
Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. On MSCOCO, it outperforms ResNe...
We present MVSNeRF, a novel neural rendering approach that can efficiently reconstruct neural radiance fields for view synthesis. Our approach can generalize across scenes (even indoor scenes, complet...
Coarse-to-fine strategies have been extensively used for the architecture design of single image deblurring networks. Extensive experiments on the GoPro and RealBlur datasets demonstrate that the prop...