Back to Top

Visual Document Retrieval – EVIE-8B, No OCR Needed

Visual document retrieval has always leaned on OCR. First, extract the text. Then, run retrieval on that text. This works fine for a clean invoice or a simple contract.

But once a document leans on layout — a table with merged cells, a chart with an axis label, a form with checkboxes scattered across a scanned page — plain text extraction starts losing the very structure that carries the meaning.

Tencent’s new open-source project, EVIE, tackles this problem by skipping the OCR step for retrieval entirely.

Instead of converting a page to text and then searching that text, EVIE looks at the page image directly and learns to match it to a query.

As a result, the flagship version, EVIE-8B, now tops the ViDoRe leaderboards for visual document retrieval.

The Challenge with Text-First Visual Document Retrieval

Traditional retrieval pipelines for documents typically follow the same pattern:

  • Run OCR to pull out raw text
  • Chunk that text into passages
  • Embed the passages and search over them

Unfortunately, this approach throws away a lot along the way. Once a table becomes a wall of text, it loses the relationship between a row and a column header.

Similarly, a watermark, a stamp, or a figure caption often jumbles together with the body text.

Worse, any OCR error early in the pipeline propagates straight through to retrieval — for example, a misread number in an invoice can quietly make a whole document invisible to search.

Treating the Page as the Unit of Visual Document Retrieval

EVIE takes a different approach. Building on the broader ColPali family of models, it stops treating text extraction as a prerequisite for search.

Instead, the system feeds a document page — scan, invoice, form, or slide — to the model as an image.

From there, the model produces a set of embeddings straight from that image, preserving layout, tables, charts, and small typography as part of the representation rather than discarding them before search even starts.

As a result, matching a query to a page comes down to comparing embeddings, not comparing strings of OCR’d text.

Visual-document-retrieval

Image source: huggingface.co

How EVIE Scores Relevance: Late Interaction with MaxSim

To score relevance, EVIE uses the same ColBERT-style late interaction mechanism that made ColPali effective for this problem.

Rather than compressing an entire page into one single vector, the model keeps one embedding per token on the query side and per visual patch on the document side. It then computes relevance between a query and a page as:

S(Q, D) = Σ max(qᵢ · dⱼ)

In other words, for every query token, the model finds its best-matching patch on the page and sums those best matches. Researchers call this technique MaxSim.

Although it costs more to compute than comparing two single vectors, it lets the model perform fine-grained matching — so a query about “the total in the far-right column” can actually land on that specific region of the page instead of getting averaged away into a single page-level summary.

Architecture Behind EVIE’s Visual Document Retrieval

EVIE’s backbone, ColQwen3.5, pairs a Qwen3.5 vision-language model with a ColPali-style late-interaction head. The family includes:

  • EVIE-8B, the flagship “teacher” model. It builds on Qwen3.5-9B, runs at roughly 8.41B parameters, and produces 4096-dimensional token embeddings.
  • EVIE-4.5B, the distilled “student” model. It builds on Qwen3.5-4B at roughly 4.61B parameters, and uses a technique the team calls ARD (Anchor-preserving Relation Distillation) to distill the larger 8B teacher’s token-relation structure into the smaller model.
  • Prefix-MRL, a feature in EVIE-4.5B that lets a single 2048-dimensional projection head shrink at inference time to 64, 128, 256, 512, or 1024 dimensions — no retraining or model-swapping required. In exchange for a small accuracy hit, teams get a much smaller index.
  • HAC (Hierarchical Agglomerative Clustering), a training-free compression step for large-scale deployment. It reduces the roughly 750 visual patch vectors per page down to as few as 32, which shrinks a million-page index to about 3.81 GiB.

Results

Across the ViDoRe benchmarks — the standard testbed for visual document retrieval, now spanning V1, V2, and V3 plus the multilingual JinaVDR suite — EVIE-8B currently ranks first:

ModelParamsEmbed DimViDoRe V1ViDoRe V2ViDoRe V3
EVIE-8B8.41B4096D92.1874.2366.75
EVIE-4.5B4.61B64–2048D92.0773.3866.02
webAI-ColVec1.1-8b8.40B640D91.3065.8265.32
nemotron-colembed-vl-8b-v28.80B4096D92.6565.1663.54

Notably, the smaller EVIE-4.5B model comes within a point of the flagship’s accuracy on most benchmarks, even though the team distilled it from the 8B teacher.

Consequently, it offers a useful trade-off for teams that need retrieval at scale rather than the absolute top score.

Broader Implications

The idea underneath EVIE — represent a document by its visual structure and score relevance with fine-grained, patch-level matching rather than a single pooled vector — extends well beyond plain document search. For instance, it can improve:

  • Retrieval-augmented generation (RAG) over scanned or visually complex documents, where OCR errors have historically hurt downstream answer quality
  • Enterprise search over invoices, forms, and reports, where layout carries meaning that plain text loses
  • Multilingual document retrieval, since the model works from the page image and query text jointly instead of depending on a language-specific OCR pipeline

Ultimately, this reflects a wider shift in document AI.

Rather than forcing every document into a single text-only representation, more models now work with the document as it actually appears — tables, charts, layout and all.

Conclusion

EVIE-8B shows that visual document retrieval doesn’t have to trade accuracy for going OCR-free.

By combining a strong vision-language backbone with ColBERT-style late interaction, plus practical compression tricks like Prefix-MRL and HAC, Tencent’s EVIE family delivers state-of-the-art results on ViDoRe while staying light enough to index and serve at scale.

The model weights, training pipelines, and evaluation code are open-sourced on GitHub (Tencent/EVIE) and Hugging Face (tencent/EVIE-8B) for the community to explore and build on.

. . .

Leave a Comment

Your email address will not be published. Required fields are marked*


Be the first to comment.

Back to Top

Message Sent!

If you have more details or questions, you can reply to the received confirmation email.

Back to Home