You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Two-sided patch selector: action-boundary detection on Assembly101 (C10119)

Class-agnostic action-boundary detectors: a V-JEPA 2.1 ViT-L video encoder, fine-tuned end to end, with a two-sided per-patch selector head trained with a boundary loss only. Every patch of every 0.25 s point attends to the patches of the 16 points before it (left selector) and after it (right selector), compares the two summaries per patch, and the per-patch comparisons are pooled into a boundary logit and offset. Single static view C10119_rgb of Assembly101, official splits (sequences: train 393 / validation 120 / test 167).

Status: research checkpoints accompanying work in progress (one seed per run). Citation: to be added.

Contents

Each folder holds one run:

  • ckpt_latest.pth: the final state {"epoch": 14, "ema_head": ..., "ema_encoder": ..., "slim": True} (EMA weights after epoch 15 of 15; no checkpoint selection).
  • config.json: the full resolved configuration and the command-line overrides. Paths are relative to the code repository; <frame cache root> is where the Assembly101 frame cache was placed.
  • results.json: final metrics, with decoding thresholds chosen on validation and test scored at them.
Folder Targets Training input Test F1, fine events (±0.25 / 0.5 / 1 / 2 s) Test F1, coarse boundaries
split1_patch_late_ld09_fine fine events 1,024-frame crops 69.2 / 75.7 / 78.7 / 81.5 14.9 / 19.6 / 22.9 / 25.1 (fires on every fine event)
split1_patch_late_ld09_coarse coarse segment boundaries 1,024-frame crops 34.3 / 41.2 / 45.4 / 47.9 43.2 / 49.6 / 56.2 / 62.1
split1_patch_late_ld09_latent_whole fine + coarse, latent-boundary likelihood (noisy-OR) whole sequences 70.4 / 76.8 / 79.6 / 82.8 39.3 / 43.9 / 50.6 / 56.5
split1_patch_late_ld09_twohead_whole fine + coarse, independent heads (control) whole sequences 70.7 / 76.8 / 79.6 / 82.7 39.3 / 45.8 / 51.0 / 57.0

F1 uses one-to-one matching at each tolerance. Fine events are every fine-segment start and end of Assembly101's fine-grained annotations (events less than 0.5 s from a sequence end dropped, events within 0.25 s merged); coarse boundaries are the changes between coarse segments.

Shared recipe: whole encoder fine-tuned from the released V-JEPA 2.1 ViT-L (distilled; run at 256 px), lr 2e-5 at the top with layer-wise decay 0.9, ±3-tubelet attention band, two interleaved tubelet phases (one point per 0.25 s), selector W = 16 points per side, width 256, 4 heads, head lr 2e-4, EMA 0.999, 15 epochs, one seed.

Loading

Requires the project code (packages tas and boundary_segmentation) and the base V-JEPA 2.1 ViT-L weights (vjepa2_1_vitl_dist_vitG_384.pt, at the path in config.json): ema_encoder holds the encoder's trainable tensors, applied on top of the base model.

import json, torch
from boundary_segmentation import train as bt
from boundary_segmentation.config import build
from tas.run import encoder_keys

run = "split1_patch_late_ld09_fine"
c = json.load(open(f"{run}/config.json"))
overrides = [o for o in c["overrides"] if not o.startswith("data.cache_dir=")] + ["data.cache_dir=<frame cache root>"]
cfg = build(c["dataset"], c["arm"], c["split"], overrides)
encoder = bt.build_encoder(cfg.exp).eval()
head = bt.build_head(cfg.boundary, encoder).eval()
ck = torch.load(f"{run}/ckpt_latest.pth", map_location="cpu", weights_only=False)
encoder.load_trainable_state_dict(ck[encoder_keys(encoder)[1]])
head.load_state_dict(ck["ema_head"])

Licence and attribution

These weights are released under CC BY-NC 4.0 (non-commercial use only).

Neither the Assembly101 authors nor Meta endorse this work.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Moose27/assembly101-boundary-selector