|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | YOLOP:You Only Look Once for Panoptic Driving Perception显示文摘A panoptic driving perception system is an essential part of autonomous driving.A high-precision and real-time perception system can assist the vehicle in making reasonable decisions while driving.We present a panoptic driving perception network(you only look once for panoptic(YOLOP))to perform traffic object detection,drivable area segmentation,and lane detection simultaneously.It is composed of one encoder for feature extraction and three decoders to handle the specific tasks.Our model performs extremely well on the challenging BDD100K dataset,achieving state-of-the-art on all three tasks in terms of accuracy and speed.Besides,we verify the effectiveness of our multi-task learning model for joint training via ablative studies.To our best knowledge,this is the first work that can process these three visual perception tasks simultaneously in real-time on an embedded device Jetson TX2(23 FPS),and maintain excellent accuracy.To facilitate further research,the source codes and pre-trained models are released at http://gffzz188fe103f8f1460aswwb0wfvf6qk06uu5.ffgz.tsg.suse.edu.cn/hustvl/YOLOP. | Dong Wu Man-Wen Liao Wei-Tian Zhang Xing-Gang Wang Xiang Bai Wen-Qing Cheng Wen-Yu Liu | 2022 | Machine Intelligence Research2022,19,6: | 8 |
| 2 | Multi-dimensional Classification via Selective Feature Augmentation显示文摘In multi-dimensional classification(MDC), the semantics of objects are characterized by multiple class spaces from different dimensions. Most MDC approaches try to explicitly model the dependencies among class spaces in output space. In contrast, the recently proposed feature augmentation strategy, which aims at manipulating feature space, has also been shown to be an effective solution for MDC. However, existing feature augmentation approaches only focus on designing holistic augmented features to be appended with the original features, while better generalization performance could be achieved by exploiting multiple kinds of augmented features.In this paper, we propose the selective feature augmentation strategy that focuses on synergizing multiple kinds of augmented features.Specifically, by assuming that only part of the augmented features is pertinent and useful for each dimension′s model induction, we derive a classification model which can fully utilize the original features while conduct feature selection for the augmented features. To validate the effectiveness of the proposed strategy, we generate three kinds of simple augmented features based on standard k NN, weighted k NN, and maximum margin techniques, respectively. Comparative studies show that the proposed strategy achieves superior performance against both state-of-the-art MDC approaches and its degenerated versions with either kind of augmented features. | Bin-Bin Jia Min-Ling Zhang | 2022 | Machine Intelligence Research2022,19,1: | 5 |
| 3 | Paradigm Shift in Natural Language Processing显示文摘In the era of deep learning, modeling for most natural language processing (NLP) tasks has converged into several mainstream paradigms. For example, we usually adopt the sequence labeling paradigm to solve a bundle of tasks such as POS-tagging, named entity recognition (NER), and chunking, and adopt the classification paradigm to solve tasks like sentiment analysis. With the rapid progress of pre-trained language models, recent years have witnessed a rising trend of paradigm shift, which is solving one NLP task in a new paradigm by reformulating the task. The paradigm shift has achieved great success on many tasks and is becoming a promising way to improve model performance. Moreover, some of these paradigms have shown great potential to unify a large number of NLP tasks, making it possible to build a single model to handle diverse tasks. In this paper, we review such phenomenon of paradigm shifts in recent years, highlighting several paradigms that have the potential to solve different NLP tasks. | Tian-Xiang Sun Xiang-Yang Liu Xi-Peng Qiu Xuan-Jing Huang | 2022 | Machine Intelligence Research2022,19,3: | 4 |
| 4 | Large-scale Multi-modal Pre-trained Models: A Comprehensive Survey显示文摘With the urgent demand for generalized deep models,many pre-trained big models are proposed,such as bidirectional encoder representations(BERT),vision transformer(ViT),generative pre-trained transformers(GPT),etc.Inspired by the success of these models in single domains(like computer vision and natural language processing),the multi-modal pre-trained big models have also drawn more and more attention in recent years.In this work,we give a comprehensive survey of these models and hope this paper could provide new insights and helps fresh researchers to track the most cutting-edge works.Specifically,we firstly introduce the background of multi-modal pre-training by reviewing the conventional deep learning,pre-training works in natural language process,computer vision,and speech.Then,we introduce the task definition,key challenges,and advantages of multi-modal pre-training models(MM-PTMs),and discuss the MM-PTMs with a focus on data,objectives,network architectures,and knowledge enhanced pre-training.After that,we introduce the downstream tasks used for the validation of large-scale MM-PTMs,including generative,classification,and regression tasks.We also give visualization and analysis of the model parameters and results on representative downstream tasks.Finally,we point out possible research directions for this topic that may benefit future works.In addition,we maintain a continuously updated paper list for large-scale pre-trained multi-modal big models:http://gffzz188fe103f8f1460aswwb0wfvf6qk06uu5.ffgz.tsg.suse.edu.cn/wangxiao5791509/MultiModal_BigModels_Survey. | Xiao Wang Guangyao Chen Guangwu Qian Pengcheng Gao Xiao-Yong Wei Yaowei Wang Yonghong Tian Wen Gao | 2023 | Machine Intelligence Research2023,20,4: | 3 |
| 5 | A Review of Predictive and Contrastive Self-supervised Learning for Medical Images显示文摘Over the last decade, supervised deep learning on manually annotated big data has been progressing significantly on computer vision tasks. But, the application of deep learning in medical image analysis is limited by the scarcity of high-quality annotated medical imaging data. An emerging solution is self-supervised learning (SSL), among which contrastive SSL is the most successful approach to rivalling or outperforming supervised learning. This review investigates several state-of-the-art contrastive SSL algorithms originally on natural images as well as their adaptations for medical images, and concludes by discussing recent advances, current limitations, and future directions in applying contrastive SSL in the medical domain. | Wei-Chien Wang Euijoon Ahn Dagan Feng Jinman Kim | 2023 | Machine Intelligence Research2023,20,4: | 2 |
| 6 | Darboux frames snakes and super-quadrics: geometry from the bottom up显示文摘 | Ferrie F P Lagarde J Whaite P | 1993 | IEEE Transactions on Pattern Analysis and Machine Intelligence1993,15,8: | 2 |
| 7 | An Unbiased Detector of Curvilinear Structures 显示文摘 | Steger C | 1998 | IEEE Transactions on Pattern Analysis and Machine Intelligence1998,20,2: | 2 |
| 8 | Classification using adaptive wavelets for feature extraction显示文摘 | Yvette M Danny C Jerry K | 1997 | IEEE Transactions on Pattern Analysis and Machine Intelligence1997,19,10: | 2 |
| 9 | Matching perspective views of a polyhedron using circuits显示文摘 | Gu WK Yang JY Huang TS | 1987 | IEEE Transaction on Pattern Analysis and Machine Intelligence1987,9,3: | 2 |
| 10 | A Handwritten Character Recognition System Using Directional Element Feature and Asymmetric Mahalanobis Distance显示文摘 | Kato N Suzuki M Omachi S | 1999 | IEEE Transactions on Pattern Analysis and Machine Intelligence1999,21,3: | 2 |
| 11 | Analysis of thinning algorithms using mathematical morphology 显示文摘 | JANG B K CHIN R T | 1990 | Pattern Analysis and Machine Intelligence1990,12,6: | 2 |
| 12 | Correction to:YOLOP:You Only Look Once for Panoptic Driving Perception显示文摘Correction to:YOLOP:You Only Look Once for Panoptic Driving Perception DOI:10.1007/s11633-022-1339-y Authors:Dong Wu,Man-Wen Liao,Wei-Tian Zhang,Xing-Gang Wang,Xiang Bai,Wen-Qing Cheng,Wen-Yu Liu The article YOLOP:You Only Look Once for Panoptic Driving Perception,written by Dong Wu,Man-Wen Liao,Wei-Tian Zhang,Xing-Gang Wang,Xiang Bai,Wen-Qing Cheng,Wen-Yu Liu,was originally published without Open Access. | Dong Wu Man-Wen Liao Wei-Tian Zhang Xing-Gang Wang Xiang Bai Wen-Qing Cheng Wen-Yu Liu | 2023 | Machine Intelligence Research2023,20,6: | 2 |
| 13 | Robust face recognition via sparse representation显示文摘 | Wright J Yang A Y Ganesh A | | Pattern Analysis Machine Intelligence0,,: | 2 |
| 14 | HMM-Based On-line Handwriting Recognition 显示文摘 | Hu J Brown MK Turin W | 1996 | IEEE Transactions on Pattern Analysis and Machine Intelligence1996,18,10: | 2 |
| 15 | VLP:A Survey on Vision-language Pre-training显示文摘In the past few years,the emergence of pre-training models has brought uni-modal fields such as computer vision(CV)and natural language processing(NLP)to a new era.Substantial works have shown that they are beneficial for downstream uni-modal tasks and avoid training a new model from scratch.So can such pre-trained models be applied to multi-modal tasks?Researchers have ex-plored this problem and made significant progress.This paper surveys recent advances and new frontiers in vision-language pre-training(VLP),including image-text and video-text pre-training.To give readers a better overall grasp of VLP,we first review its recent ad-vances in five aspects:feature extraction,model architecture,pre-training objectives,pre-training datasets,and downstream tasks.Then,we summarize the specific VLP models in detail.Finally,we discuss the new frontiers in VLP.To the best of our knowledge,this is the first survey focused on VLP.We hope that this survey can shed light on future research in the VLP field. | Fei-Long Chen Du-Zhen Zhang Ming-Lun Han Xiu-Yi Chen Jing Shi Shuang Xu Bo Xu | 2023 | Machine Intelligence Research2023,20,1: | 2 |
| 16 | Log-polar wavelet energy signatures for rotation and scale invariant texture classification 显示文摘 | Pun Chi-Man Lee Moon-Chuen | 2003 | IEEE Transactions on Pattern Analysis and Machine Intelligence2003,25,5: | 2 |
| 17 | Fractal-based description of nature scenes显示文摘 | Pentland A P | 1984 | IEEE Trans on Pattern Analysis and Machine Intelligence1984,6,6: | 2 |
| 18 | Multichannel texture analysis using localized spatial filters显示文摘 | Bovic A C Clark C Geisler W S | 1990 | IEEE Transactions on Pattern Analysis and Machine Intelligence1990,12,1: | 2 |
| 19 | validity measure for Fuzzy Clustering显示文摘 | XIE X BENI G A | 1991 | IEEE Transactions on Pattern Analysis and Machine Intelligence1991,13,8: | 2 |
| 20 | Large vocabulary recognition of on - line handwritten cursive words显示文摘 | SENI G SRIHARI R K NASRABADI N | 1996 | IEEE Transactions on Pattern Analysis and Machine Intelligence1996,18,7: | 2 |