V-JEPA2 ViT-L (MLX) — encoder + JEPA predictor

Apple MLX fp16 port of Meta's V-JEPA2 ViT-L/16 (facebook/vjepa2-vitl-fpc64-256): video/image embedding extraction plus the JEPA latent-space predictor (masked world-model). MIT.

pip install vjepa2-mlx   # https://github.com/xocialize/vjepa2-mlx
vjepa2-mlx -i clip.mp4 --task embed -o emb.npy
from vjepa2_mlx.pipeline_mlx import embed_video
emb = embed_video("clip.mp4", num_frames=16)   # (1024,)
  • Arch: ViT-L/16, 24 layers, hidden 1024, 16 heads, 3D-tubelet Conv3d patch embed, 3D-RoPE; predictor 384/12/12.
  • Parity vs PyTorch (cpu fp32): encoder rel 2.66e-5 · predictor rel 1.67e-6.
  • Precision: fp16 (~650 MB; encoder fp16 rel 4.7e-3).

MIT (© Meta Platforms). Action-conditioned (robotics) predictor not included — separate ViT-g model.

Downloads last month
19
Safetensors
Model size
0.3B params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-community/V-JEPA2-vitl-fpc64-256