Temporal Localization

Temporal localization marks when an event starts and ends in an untrimmed video. Temporal action recognition may say “goal attempt”; localization must output a segment such as 12.0s to 22.0s. This is the bridge from clip classification to usable timelines, alerts, and search indexes.

Temporal IoU and mAP

For a predicted segment and ground truth , temporal intersection-over-union is

Detectors score candidate segments, then evaluate mean average precision at one or more tIoU thresholds. Sliding-window inference is the simplest proposal mechanism; modern video transformers can predict sparse temporal proposals directly.

Worked example

For ground truth , compare two proposed segments:

segmentintervalintersectionunion spantIoUboundary error
good proposalstart s, end s
bad proposalno overlap

The good segment overlaps substantially but starts two seconds early and ends one second early. For a highlight reel that may be fine; for a safety trigger it may not.

Temporal localization compares predicted and true action intervals by their overlap and union span.

Caveats

Ambiguous boundaries make annotation and evaluation noisy. High tIoU thresholds punish small timing errors on short actions. Models trained on trimmed clips often fail on background-heavy video because the negative temporal context is different.

References