multilingual-e5-small β€” ExecuTorch

The smallest of the multilingual E5 family. Text in, one 384-dimensional vector out, for search and retrieval that never leaves the device.

  • Source: intfloat/multilingual-e5-small β€” 12 layers, 384 dimensions, 250,037 vocabulary
  • License: mit
  • Input: input_ids and attention_mask, both [1, 256] int64
  • Output: [1, 384], mean-pooled and L2-normalised inside the graph

The recipe is in the graph, and it was read off this repo

sentence-transformers stores it per model, and the shelf's seven embedding models do not agree. This one pools mean and normalises, read from 1_Pooling/config.json and modules.json rather than inferred from the family name. Getting it wrong does not throw; it returns vectors that look fine and rank wrong.

The prefix is not in the graph

This model is trained with query: in front of the text and expects it at inference. That happens before tokenisation, so the .pte never sees it as anything but tokens β€” and leaving it out does not throw. It returns a plausible vector that retrieves worse.

Verification

build file size (MB) Mac ms* worst cosine vs eager retrieval budget
fp32 embed_multilingual_e5_small_xnnpack_fp32.pte 470.2 29.3 1.000000 0%
fp16 embed_multilingual_e5_small_xnnpack_fp16.pte 235.3 52.2 1.000000 1%
Core ML (fp16, iOS) embed_multilingual_e5_small_coreml_all.pte 235.5 4.3 0.999979 15%

*Mac arm64, median of 10, one 256-token sequence β€” a reference point for relative cost, not a device number. Torch eager fp32 on the same machine is 18.0 ms.

Cosine is measured against the model run in eager through its own pooling, over eight sentences. The last column is the one that decides: rank those eight against each other, and ask whether this build's score error is smaller than the gap between the document a query retrieves and the runner-up. Every shipped build keeps all eight top-1 results.

Not shipped: int8

embed_multilingual_e5_small_xnnpack_int8.pte is 406.7 MB against fp16's 235.3 MB. Dynamic int8 quantises the linear weights and leaves the token embedding table in fp32, and here that table is 384 MB of the 470.2 MB model β€” 82%. The size a build comes out at is 0.5 + 1.5 x (table share) times the fp16 build; at 82% that is 1.73, so there was never a smaller file to be had.

It is withheld on the number that decides. Ranking the eight test sentences against each other, this build moves a pair score by at most 0.0037 while the closest fp32 decision β€” the gap between the document a query retrieves and the runner-up β€” is 0.0105. That is 35% of the room available, against a bar of 50%.

Correlation reads 0.999568 for this build, which no correlation gate would stop.

torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlboydaisuke/multilingual-e5-small-ExecuTorch

Quantized
(270)
this model