Two-sided patch selector: action-boundary detection on Assembly101 (C10119)
Class-agnostic action-boundary detectors: a V-JEPA 2.1 ViT-L video encoder, fine-tuned end to end, with a two-sided per-patch selector head trained with a boundary loss only. Every patch of every 0.25 s point attends to the patches of the 16 points before it (left selector) and after it (right selector), compares the two summaries per patch, and the per-patch comparisons are pooled into a boundary logit and offset. Single static view C10119_rgb of Assembly101, official splits (sequences: train 393 / validation 120 / test 167).
Status: research checkpoints accompanying work in progress (one seed per run). Citation: to be added.
Contents
Each folder holds one run:
ckpt_latest.pth: the final state{"epoch": 14, "ema_head": ..., "ema_encoder": ..., "slim": True}(EMA weights after epoch 15 of 15; no checkpoint selection).config.json: the full resolved configuration and the command-line overrides. Paths are relative to the code repository;<frame cache root>is where the Assembly101 frame cache was placed.results.json: final metrics, with decoding thresholds chosen on validation and test scored at them.
| Folder | Targets | Training input | Test F1, fine events (±0.25 / 0.5 / 1 / 2 s) | Test F1, coarse boundaries |
|---|---|---|---|---|
split1_patch_late_ld09_fine |
fine events | 1,024-frame crops | 69.2 / 75.7 / 78.7 / 81.5 | 14.9 / 19.6 / 22.9 / 25.1 (fires on every fine event) |
split1_patch_late_ld09_coarse |
coarse segment boundaries | 1,024-frame crops | 34.3 / 41.2 / 45.4 / 47.9 | 43.2 / 49.6 / 56.2 / 62.1 |
split1_patch_late_ld09_latent_whole |
fine + coarse, latent-boundary likelihood (noisy-OR) | whole sequences | 70.4 / 76.8 / 79.6 / 82.8 | 39.3 / 43.9 / 50.6 / 56.5 |
split1_patch_late_ld09_twohead_whole |
fine + coarse, independent heads (control) | whole sequences | 70.7 / 76.8 / 79.6 / 82.7 | 39.3 / 45.8 / 51.0 / 57.0 |
F1 uses one-to-one matching at each tolerance. Fine events are every fine-segment start and end of Assembly101's fine-grained annotations (events less than 0.5 s from a sequence end dropped, events within 0.25 s merged); coarse boundaries are the changes between coarse segments.
Shared recipe: whole encoder fine-tuned from the released V-JEPA 2.1 ViT-L (distilled; run at 256 px), lr 2e-5 at the top with layer-wise decay 0.9, ±3-tubelet attention band, two interleaved tubelet phases (one point per 0.25 s), selector W = 16 points per side, width 256, 4 heads, head lr 2e-4, EMA 0.999, 15 epochs, one seed.
Loading
Requires the project code (packages tas and boundary_segmentation) and the base V-JEPA 2.1 ViT-L weights
(vjepa2_1_vitl_dist_vitG_384.pt, at the path in config.json): ema_encoder holds the encoder's trainable tensors,
applied on top of the base model.
import json, torch
from boundary_segmentation import train as bt
from boundary_segmentation.config import build
from tas.run import encoder_keys
run = "split1_patch_late_ld09_fine"
c = json.load(open(f"{run}/config.json"))
overrides = [o for o in c["overrides"] if not o.startswith("data.cache_dir=")] + ["data.cache_dir=<frame cache root>"]
cfg = build(c["dataset"], c["arm"], c["split"], overrides)
encoder = bt.build_encoder(cfg.exp).eval()
head = bt.build_head(cfg.boundary, encoder).eval()
ck = torch.load(f"{run}/ckpt_latest.pth", map_location="cpu", weights_only=False)
encoder.load_trainable_state_dict(ck[encoder_keys(encoder)[1]])
head.load_state_dict(ck["ema_head"])
Licence and attribution
These weights are released under CC BY-NC 4.0 (non-commercial use only).
- Training data: Assembly101 (Sener et al., CVPR 2022; https://assembly-101.github.io), licensed under CC BY-NC 4.0 (https://creativecommons.org/licenses/by-nc/4.0/). Changes: the models were trained on the C10119 static view's frames (sampled at 4 fps, 256 px) with boundary targets derived from its coarse and fine annotations.
- Base model: V-JEPA 2.1 ViT-L by Meta FAIR (arXiv 2603.14482; https://github.com/facebookresearch/vjepa2, branch
vjepa2_1). Changes: the whole encoder was fine-tuned. Use is also subject to the V-JEPA 2.1 licence.
Neither the Assembly101 authors nor Meta endorse this work.