--- base_model: Qwen/Qwen3.5-4B library_name: peft license: apache-2.0 pipeline_tag: image-text-to-text tags: - lora - peft - egocentric-video - proactive-assistant - streaming-video - eccv-2026 - wearable-ai - egocentric-vision --- # EgoProactive 4B — proactive assistant verbalizer (LoRA) 🏆 LoRA adapter for **Qwen/Qwen3.5-4B**, from Team Ambient's entry to the **EgoProactive** track of the [Wearable AI Challenge @ ECCV 2026](https://wearable-ai-workshop.github.io/) — **1st place, large division**. | | | |---|---| | Division | Large (2B+) | | Result | **0.7179 macro-F1** (1st of 7) | | Base model | [`Qwen/Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) | | Adapter | LoRA r=32, α=64, dropout 0.05, on all 7 attention/MLP projections | | Merged size | 4.5393 B parameters | ## What it does Given egocentric video arriving as ~8-second chunks plus the user's opening query, the model decides after each chunk whether to **speak** or **stay silent**. Rather than generating `$interrupt$` / `$silent$`, it emits a **single token** — `yes` or `no` — and the decision is read off those two logits: ``` p_interrupt = softmax([logit_no, logit_yes])[1] interrupt if p_interrupt >= tau ``` Operating point: **tau = 0.55**. Utterances are templated at inference; the challenge metric scores only the timing decision, not the wording. ## Usage ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor from peft import PeftModel BASE = "Qwen/Qwen3.5-4B" proc = AutoProcessor.from_pretrained(BASE) model = PeftModel.from_pretrained( AutoModelForImageTextToText.from_pretrained(BASE, dtype=torch.bfloat16, device_map="cuda"), "ambient-intelligence-labs/egoproactive-4b-lora").eval() YES = proc.tokenizer.encode("yes", add_special_tokens=False)[0] NO = proc.tokenizer.encode("no", add_special_tokens=False)[0] # inputs: cumulative frames (strided to 32) + query + last 4 dialogue turns logits = model(**inputs).logits[0, -1] p_interrupt = torch.softmax(torch.stack([logits[NO], logits[YES]]).float(), 0)[1].item() speak = p_interrupt >= 0.55 ``` The **dialogue history is required** — without it the model has no way to know it has already spoken and fires on every chunk. Frames are cumulative from the start of the video, strided to a cap of 32, resized to a 512px maximum side. ## Training Fine-tuned on the released validation videos (seen twice) plus a synthetic corpus of 234 clips annotated by a tool-calling video agent that inspects the footage before placing each cue — 13,730 rows, 46.5% interrupt. lr 1e-4 cosine, warmup 0.03, batch size 1 × 8 accumulation, bf16 with gradient checkpointing, 1 epoch. - Code: - Data: [`ambient-intelligence-labs/egoproactive-synth-annotations`](https://huggingface.co/datasets/ambient-intelligence-labs/egoproactive-synth-annotations) ## Citation ```bibtex @techreport{umapathi2026speak, title = {Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision}, author = {Umapathi, Logesh Kumar}, year = {2026}, institution = {Team Ambient}, note = {Wearable AI Challenge @ ECCV 2026, EgoProactive track} } ```