Temporal Localization
Temporal localization marks when an event starts and ends in an untrimmed video. Temporal action recognition may say “goal attempt”; localization must output a segment such as 12.0s to 22.0s. This is the bridge from clip classification to usable timelines, alerts, and search indexes.
Temporal IoU and mAP
For a predicted segment and ground truth , temporal intersection-over-union is
Detectors score candidate segments, then evaluate mean average precision at one or more tIoU thresholds. Sliding-window inference is the simplest proposal mechanism; modern video transformers can predict sparse temporal proposals directly.
Worked example
For ground truth , compare two proposed segments:
| segment | interval | intersection | union span | tIoU | boundary error |
|---|---|---|---|---|---|
| good proposal | start s, end s | ||||
| bad proposal | no overlap |
The good segment overlaps substantially but starts two seconds early and ends one second early. For a highlight reel that may be fine; for a safety trigger it may not.
Caveats
Ambiguous boundaries make annotation and evaluation noisy. High tIoU thresholds punish small timing errors on short actions. Models trained on trimmed clips often fail on background-heavy video because the negative temporal context is different.
References
- Wu et al., 2021, Towards High-Quality Temporal Action Detection with Sparse Proposals
- Kay et al., 2017, The Kinetics Human Action Video Dataset
Nav