Cold-Start Forecasting

Cold-start forecasting covers entities with little or no usable history. Examples include new products, newly instrumented machines, new regions, new categories, and series that were absent from backtesting but appear in production.

Cold starts are not rare edge cases in operational systems. They should be represented explicitly in training, validation, fallback logic, and monitoring.

Cold-start cases

A new series with no observations has metadata and future covariates but no target history. A short-history series has fewer observations than the required lag or rolling-window context. A new category contains levels not seen during training. A production-only series appears in live inference but was absent from backtesting or ensemble fitting.

These cases differ. A product with no sales because it has not launched is not the same as a mature product with zero demand.

Strategies

StrategyAssumptionWhen appropriateMain risk
Dropthe series can be ignoredoffline model comparisonunusable in production if every entity needs a forecast
Zero paddingno history means no prior activitygenuine pre-launch entitiestreats “not measured” as observed zero demand
Missing-value paddingabsence of history is informativewhen zero and unknown must differneeds models that consume missing indicators
Partial-history trainingvariable context length is learnablemany short seriesrequires masking or flexible architectures
Global-model transferrelated series share structurelarge panels with metadatanegative transfer if series are heterogeneous
Metadata analoguessimilar entities behave similarlyrich static attributeswrong analogue class misleads the forecast
Baseline fallbacka deterministic policy is good enoughlast resort for any entitytoo coarse if used where real signal exists

Drop excludes series lacking sufficient context. This is defensible for offline model comparison but can be unacceptable in production if forecasts are required for every active entity.

Zero padding fills missing history with zeros. It assumes no prior activity, which may be appropriate for some launch processes but misleading when the series existed before measurement began.

Missing-value padding preserves the distinction between no history and zero demand. Models can use missing indicators to learn how short context differs from true low demand.

Partial-history training trains models to work with variable-length context using masks, shorter lag sets, or architectures that tolerate missing context.

Global-model transfer uses patterns learned from related series. Static metadata and future-known covariates become especially important when target history is unavailable.

Metadata-based analogues infer from similar entities such as category, region, brand, lifecycle stage, capacity class, or historical launch cohort.

Baseline fallback uses deterministic policies such as global mean, category mean, seasonal baseline, partition-level best model, or a globally selected model.

Fallback policy

A fallback policy should be part of the model definition, not an ad hoc production patch. It should specify eligibility rules, priority order, deterministic tie-breaking, output units, clipping behavior, and monitoring counters.

For example, a new retail item might use category-level average demand per active store, multiplied by planned store exposure. A short-history item might use a global gradient-boosted model if enough metadata is available, otherwise a category baseline.

Evaluation

Cold-start entities should be evaluated as a separate population. A global average can hide severe cold-start errors because mature series dominate the metric. Backtesting can simulate cold starts by hiding early history, holding out recently launched entities, or evaluating series absent from the ensemble-fitting period.

Practical guidance

  • Decide whether missing pre-launch history means unknown, zero, or not applicable.
  • Train and evaluate at least one fallback that does not require target history.
  • Use static metadata, exposure, and planned covariates for no-history entities.
  • Monitor fallback frequency and forecast quality separately from regular forecasts.
  • Keep fallback behavior deterministic and auditable.

Common failure modes

  • Padding with zeros and then treating those zeros as observed demand.
  • Evaluating only mature series and discovering cold-start failures in production.
  • Letting unseen categories map silently to arbitrary encodings.
  • Using per-series model selection when a series has too few validation points.
  • Hiding fallback usage in logs instead of surfacing it as a metric.

Connections

Cold starts tie forecasting data and covariates to production fallback rules. Machine learning forecasting can borrow strength from related series, while intermittent demand and demand forecasting define common operational cases.

References