Zeva-Ego Action Encoder

Model description

This repository contains the frozen Action-Centric Encoder (ACE) used by Zeva-Ego to obtain action-token supervision from RGB visual transitions. Given two frames from the same video, the encoder produces a set of task tokens that represent the observed interaction and a smaller set of environment tokens. The representation is intended for labeling action-unlabeled egocentric video before vision-language-action model mid-training.

This is an inference release matching the checkpoint contract implemented by the Zeva repository. The checkpoint contains the complete Stage 2 model state required by its strict loader. The reconstruction decoder is included for checkpoint compatibility but is not executed by the RGB-pair token-labeling path. Optimizer and scheduler states, training counters, dataset indexes, and machine-specific configuration are not included.

Checkpoint index

Checkpoint Framework Role
Zeva-Ego π0.5 JAX 55k JAX / OpenPI Primary egocentric-video mid-training initializer
Zeva-Ego π0.5 LeRobot 40k PyTorch / LeRobot LeRobot-compatible egocentric-video mid-training initializer
Zeva-Ego π0.5 RoboTwin JAX 60k JAX / OpenPI RoboTwin downstream policy initialized from JAX 55k
Zeva-Ego Action Encoder PyTorch Frozen ACE labeler for RGB visual transitions

Model details

Property Description
Input Two RGB frames from one visual transition
Visual backbone Frozen DINOv2 ViT-B/14 with registers
Task-token output 16 × 128
Environment-token output 4 × 128
Text input None
Output semantics Learned visual-action representations, not executable robot commands
Intended role Offline pseudo-labeling for egocentric-video mid-training

The two frames must have identical spatial dimensions and must be resized to a supported DINOv2 patch-aligned resolution before encoding. We recommend a one-second interval on the temporally normalized video. If a video has been slowed down, the frame interval should be measured after that slowdown; videos used without temporal slowdown retain a one-second source-time interval.

Repository contents

  • weights/action_encoder.pt: complete Stage 2 inference parameters and portable architecture configuration;
  • manifest.json: public artifact hash and interface metadata; and
  • LICENSE and NOTICE: release terms and third-party attribution.

The DINOv2 source and pretrained backbone are resolved from the pinned upstream revision 7764ea0f912e53c92e82eb78a2a1631e92725fc8 at runtime and are not redistributed in this repository. Offline inference requires this revision and the corresponding dinov2_vitb14_reg4_pretrain.pth file in the PyTorch hub cache.

Usage

Install the pipelines/ego_action_encoder package from the Zeva feature/zeva_ego branch, then prepare a uint8 RGB array with shape [N, 2, H, W, 3] or [N, 2, 3, H, W]:

python pipelines/ego_action_encoder/scripts/encode_pairs.py \
  --input <rgb-pairs.npy> \
  --checkpoint weights/action_encoder.pt \
  --output outputs/example

The command writes:

  • task_tokens.npy with shape [N, 16, 128]; and
  • environment_tokens.npy with shape [N, 4, 128].

The encoder receives images only and does not require a language instruction.

Intended use

The model is intended for research on visual action representation learning, egocentric-video annotation, and action-token supervision for downstream policy mid-training. It may also be used to study cross-domain action geometry between human and robot videos.

Out-of-scope use

The outputs must not be interpreted as calibrated robot controls. The encoder is not a motion planner, safety controller, or zero-shot policy, and should not be used to command physical hardware directly.

Limitations

  • The representation does not identify a robot embodiment, coordinate frame, control frequency, or gripper convention.
  • Frame selection and temporal normalization materially affect the resulting token labels.
  • The frozen visual backbone may inherit limitations from its pretraining distribution.

License and attribution

Zeva-Ego encoder code and weights are released under Apache-2.0. The matching implementation is available in the Zeva repository. DINOv2 and other third-party components remain subject to their respective licenses and terms. See NOTICE for attribution. DINOv2 weights are not bundled with this release.

Citation

@misc{huang2026zevaego,
  title         = {{Zeva-Ego}: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation},
  author        = {Huang, Bingjia and Ding, Xin and Chen, Fu and Li, Kun and Sun, Wei and Wu, Hao and Liu, Yunxin and Cao, Ting},
  year          = {2026},
  eprint        = {2609.24411},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.24411}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Collection including BEN12324/zeva-ego-action-encoder

Paper for BEN12324/zeva-ego-action-encoder