UB Paderborn / Katalog / Suche / Details

Sie befinden Sich nicht im Netzwerk der Universität Paderborn. Der Zugriff auf elektronische Ressourcen ist gegebenenfalls nur via VPN oder Shibboleth (DFN-AAI) möglich. mehr Informationen...

Zur Ergebnisliste

Grounded Language-Image Pre-training

2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, p.10955-10965

2022

Details

Autor(en) / Beteiligte

Titel

Grounded Language-Image Pre-training

Ist Teil von

2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, p.10955-10965

Ort / Verlag

IEEE

Erscheinungsjahr

2022

Link zum Volltext

IEEE_Xplore

Quelle

IEEE Xplore

Beschreibungen/Notizen

This paper presents a grounded language-image pretraining (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representations semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1) When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines. 1 1 Supervised baselines on COCO object detection: Faster-RCNN w/ ResNet50 (40.2) or ResNet101 (42.0), and DyHead w/ Swin-Tiny (49.7). 2) After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3) When transferred to 13 downstream object detection tasks, a 1-shot GLIP rivals with a fully-supervised Dynamic Head. Code will be released at https://github.com/microsoft/GLIP.

Sprache: Englisch
Identifikatoren: eISSN: 2575-7075
DOI: 10.1109/CVPR52688.2022.01069
Titel-ID: cdi_ieee_primary_9879567

Format: –
Schlagworte: categorization, Computer vision, Data models, Deep learning architectures and techniques, Recognition: detection, Grounding, Head, Image recognition, Object detection, retrieval, Representation learning, Transfer/low-shot/long-tail learning, Vision + language, Visualization

Weiterführende Literatur

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX