ColPali
Safetensors
English
qwen2_vl
Edit model card

ColPali: Visual Retriever based on PaliGemma-3B with ColBERT strategy

ColQwen is a model based on a novel model architecture and training strategy based on Vision Language Models (VLMs) to efficiently index documents from their visual features. It is a Qwen2-VL-2B extension that generates ColBERT- style multi-vector representations of text and images. It was introduced in the paper ColPali: Efficient Document Retrieval with Vision Language Models and first released in this repository

This version is the untrained base version to guarantee deterministic projection layer initialization.

Usage

This version should not be used: it is solely the base version useful for deterministic LoRA initialization.

Contact

Citation

If you use any datasets or models from this organization in your research, please cite the original dataset as follows:

@misc{faysse2024colpaliefficientdocumentretrieval,
  title={ColPali: Efficient Document Retrieval with Vision Language Models}, 
  author={Manuel Faysse and Hugues Sibille and Tony Wu and Bilel Omrani and Gautier Viaud and Céline Hudelot and Pierre Colombo},
  year={2024},
  eprint={2407.01449},
  archivePrefix={arXiv},
  primaryClass={cs.IR},
  url={https://arxiv.org/abs/2407.01449}, 
}
Downloads last month
0
Safetensors
Model size
2.21B params
Tensor type
F32
·
Inference API
Unable to determine this model’s pipeline type. Check the docs .

Model tree for vidore/colqwen2-base

Finetuned
(21)
this model
Adapters
6 models
Finetunes
7 models

Space using vidore/colqwen2-base 1