Integrate with Sentence Transformers via MultiVectorEncoder

#1
by tomaarsen HF Staff - opened

Hello @xiaoxiaoshadiao and Team!

Congratulations on the big non-preview release! SOTA once again! Like the Preview model, I also wanted to integrate these 2 with Sentence Transformers to give users more inference options. I recognize that you're still actively making changes, so just let me know if you need me to incorporate changes from main or something.

Heads up, this PR was AI-generated and human-reviewed.

Pull Request overview

  • Integrate EVIE-4.5B with Sentence Transformers via MultiVectorEncoder
  • Add usage examples for full-width embeddings and smaller Matryoshka prefixes

Details

This follows the Preview integration from https://huggingface.co/tencent/EVIE-Preview-4.5B/discussions/1. I've kept the ColPali usage first, as requested there, and added the Sentence Transformers example below it.

The integration uses Transformer -> Dense -> Normalize -> MultiVectorMask, with the trained projection weights and bias extracted from this checkpoint into 1_Dense/model.safetensors. Bidirectional attention is configured through sentence_bert_config.json, and the existing 768-token visual budget is preserved. A separate retrieval chat template reproduces the ColPali input format while keeping the original chat template unchanged. The processor metadata resolves the standard Qwen3VLProcessor, so no ColPali imports or trust_remote_code are needed.

The default output is 2048-dimensional. The README also shows how to obtain smaller Matryoshka prefixes by slicing and renormalizing each token embedding.

This PR is primarily additive. The only changes to existing runtime configuration are processor_class and padding_side. The ColPali path explicitly loads ColQwen3_5Processor and hardcodes left padding in its constructor, so it does not rely on these metadata values. The existing ColPali inference code and model configuration are unchanged, and ColPali behavior should therefore remain the same.

Verified against this repository's ColPali implementation in float32 with Transformers 5.15.0: full-width token embeddings are identical, and the maximum MaxSim difference across all six supported dimensions is 1.9e-6. The reference was also checked with the pinned Transformers 5.13.1, with MaxSim differences below 3.1e-5 between versions. The README output comes from a separate default-dtype run using the public image URLs below.

To try this integration before the PR is merged, use the PR revision below.

pip install -U "sentence-transformers[image]>=6.0.0"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("tencent/EVIE-4.5B", revision="refs/pr/1")

queries = [
    "What is the variable represented on the y-axis of the graph?",
    "Total outlay is maximum in which year?",
]
documents = [
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg",
    "https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg",
]

query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# torch.Size([23, 2048]) torch.Size([755, 2048])

scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[17.0977,  7.8730,  6.9902,  4.8066],
#         [ 4.4795, 12.8516,  4.2031,  4.3779]])

Happy to adjust anything. Please let me know if you have any questions or feedback!

  • Tom Aarsen
tomaarsen changed pull request status to open

Oh, also, I was wondering if the Prefix-MRL truncation is the same as with dense models, i.e. just taking the first {64, 128, 256, 512, 1024, 2048} dimensions. Sentence Transformers currently doesn't support that out of the box like it does with dense models via a truncate_dim, but I've been considering adding it.

Sentence Transformers does support Hieracrhical Clustering out of the box, e.g.

token_pooling = HierarchicalTokenPooling(pool_factor=4)
pooled_embeddings = model.encode_document(..., token_pooling=token_pooling)
  • Tom Aarsen

Thank you for the invitation and I am currently making revisions and conducting tests.
我本地先测试一下这个库的效果 确定没问题之后 我会直接传上去的

Hi Tom — thanks for the integration and the verification.

We applied the Sentence Transformers integration on main ourselves. One change from this PR: the default visual budget for this release is 1024, not 768 (768 was Preview). processor_config longest_edge is now 1048576.

Prefix-MRL truncation is the first {64, 128, 256, 512, 1024, 2048} dimensions, then renormalize each token.

Please close this PR.

tomaarsen changed pull request status to closed
Tencent org
•
This comment has been hidden (marked as Resolved)

Thanks for applying, and answering my question. I've closed the PR 🤗

  • Tom Aarsen
Tencent org
•
This comment has been hidden (marked as Off-Topic)

你提到的 HierarchicalTokenPooling 和我们的 HAC 不是同一套。ST 这个是只按 embedding 余弦做 Ward,用 pool_factor 按比例少留 token。我们的 HAC 是在语义特征和 2D 空间位置的联合空间里做 Ward,压到固定 token 数(默认 32 个/页)。两边的聚类空间、目标个数、实现都不同,不能互相替代

Oh, thanks for explaining. That is indeed different than what Sentence Transformers does currently.
Will you also have a look at the pull request for the 8B model? Thanks in advance.
https://huggingface.co/tencent/EVIE-8B/discussions/1

  • Tom Aarsen

8B已经更新了

xiaoxiaoshadiao changed pull request status to open

My translation might be wrong, apologies if so, but I think the 8B model was not updated. I don't see the new files like modules.json.

刚刚更新了哈哈 我这边可能有一些延迟 现在已经可以了

Clipboard_Screenshot_1788787515

Haha perfect, sounds good! Thanks, I think this is very nice, and I think both are done now.

  • Tom Aarsen
tomaarsen changed pull request status to closed

Sign up or log in to comment