Zeva-Ego Action Encoder
Model description
This repository contains the frozen Action-Centric Encoder (ACE) used by Zeva-Ego to obtain action-token supervision from RGB visual transitions. Given two frames from the same video, the encoder produces a set of task tokens that represent the observed interaction and a smaller set of environment tokens. The representation is intended for labeling action-unlabeled egocentric video before vision-language-action model mid-training.
This is an inference release matching the checkpoint contract implemented by the Zeva repository. The checkpoint contains the complete Stage 2 model state required by its strict loader. The reconstruction decoder is included for checkpoint compatibility but is not executed by the RGB-pair token-labeling path. Optimizer and scheduler states, training counters, dataset indexes, and machine-specific configuration are not included.
Checkpoint index
| Checkpoint | Framework | Role |
|---|---|---|
| Zeva-Ego π0.5 JAX 55k | JAX / OpenPI | Primary egocentric-video mid-training initializer |
| Zeva-Ego π0.5 LeRobot 40k | PyTorch / LeRobot | LeRobot-compatible egocentric-video mid-training initializer |
| Zeva-Ego π0.5 RoboTwin JAX 60k | JAX / OpenPI | RoboTwin downstream policy initialized from JAX 55k |
| Zeva-Ego Action Encoder | PyTorch | Frozen ACE labeler for RGB visual transitions |
Model details
| Property | Description |
|---|---|
| Input | Two RGB frames from one visual transition |
| Visual backbone | Frozen DINOv2 ViT-B/14 with registers |
| Task-token output | 16 × 128 |
| Environment-token output | 4 × 128 |
| Text input | None |
| Output semantics | Learned visual-action representations, not executable robot commands |
| Intended role | Offline pseudo-labeling for egocentric-video mid-training |
The two frames must have identical spatial dimensions and must be resized to a supported DINOv2 patch-aligned resolution before encoding. We recommend a one-second interval on the temporally normalized video. If a video has been slowed down, the frame interval should be measured after that slowdown; videos used without temporal slowdown retain a one-second source-time interval.
Repository contents
weights/action_encoder.pt: complete Stage 2 inference parameters and portable architecture configuration;manifest.json: public artifact hash and interface metadata; andLICENSEandNOTICE: release terms and third-party attribution.
The DINOv2 source and pretrained backbone are resolved from the pinned upstream
revision 7764ea0f912e53c92e82eb78a2a1631e92725fc8 at runtime and are not
redistributed in this repository. Offline inference requires this revision and
the corresponding dinov2_vitb14_reg4_pretrain.pth file in the PyTorch hub
cache.
Usage
Install the pipelines/ego_action_encoder package from the Zeva
feature/zeva_ego branch, then prepare a uint8 RGB array with shape
[N, 2, H, W, 3] or [N, 2, 3, H, W]:
python pipelines/ego_action_encoder/scripts/encode_pairs.py \
--input <rgb-pairs.npy> \
--checkpoint weights/action_encoder.pt \
--output outputs/example
The command writes:
task_tokens.npywith shape[N, 16, 128]; andenvironment_tokens.npywith shape[N, 4, 128].
The encoder receives images only and does not require a language instruction.
Intended use
The model is intended for research on visual action representation learning, egocentric-video annotation, and action-token supervision for downstream policy mid-training. It may also be used to study cross-domain action geometry between human and robot videos.
Out-of-scope use
The outputs must not be interpreted as calibrated robot controls. The encoder is not a motion planner, safety controller, or zero-shot policy, and should not be used to command physical hardware directly.
Limitations
- The representation does not identify a robot embodiment, coordinate frame, control frequency, or gripper convention.
- Frame selection and temporal normalization materially affect the resulting token labels.
- The frozen visual backbone may inherit limitations from its pretraining distribution.
License and attribution
Zeva-Ego encoder code and weights are released under Apache-2.0. The matching
implementation is available in the
Zeva repository.
DINOv2 and
other third-party components remain subject to their respective licenses and
terms. See NOTICE for attribution. DINOv2 weights are not bundled with this
release.
Citation
@misc{huang2026zevaego,
title = {{Zeva-Ego}: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation},
author = {Huang, Bingjia and Ding, Xin and Chen, Fu and Li, Kun and Sun, Wei and Wu, Hao and Liu, Yunxin and Cao, Ting},
year = {2026},
eprint = {2609.24411},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.24411}
}