Applied time-series forecasting

Create a Predictive Sales Model

Example: predict daily units for coffee or tea from trend, weekly seasonality, and a planned promotion. Validate against future dates before producing a forecast.

Learning guide only. The dataset is synthetic; this page creates no Google Cloud resources, model calls, or charges.

INTERACTIVE WORKED EXAMPLE

Chronological train → holdout → future forecast

MODEL MAE

2.9

units / holdout day

SEASONAL BASELINE MAE

10.7

repeat last observed week

MODEL WAPE

2.0%

14-day chronological holdout

Synthetic daily sales forecastTRAIN · 106 DAYSHOLDOUTFUTURE
— actual synthetic units— unseen holdout prediction— future scenario

The future promotion value is a scenario input. It must be known or planned at forecast time. This toy regression does not estimate promotion causality, prediction intervals, stockouts, price elasticity, or uncertainty.

Choose the Google Cloud pathway

BEST FIT

Daily product sales already stored in BigQuery.

Aggregate one row per date and product, train ARIMA_PLUS, then request forecasts with ML.FORECAST. Use ARIMA_PLUS_XREG when adding external regressors and supply future values for those features. AI.FORECAST offers a TimesFM-based alternative without creating a model.

Check supported regions, limits, history requirements, and pricing. Holiday configuration is not a guarantee of correct holiday effects; US below is an example, not your business location.

Official documentation

Implementation examples

Mathematical formulas

Trend + weekly seasonality + promotion
y^t=max⁡ ⁣(0,β0+β1t/100+β2sin⁡(2πt/7)+β3cos⁡(2πt/7)+β4pt)\hat y_t=\max\!\left(0,\beta_0+\beta_1 t/100+\beta_2\sin(2\pi t/7)+\beta_3\cos(2\pi t/7)+\beta_4 p_t\right)

Ordinary least squares with engineered time features; this browser example is not ARIMA_PLUS or a Vertex AI model. Promotion effects are associations, not causal estimates.

Chronological holdout error
MAE=1H∑t=1H∣yt−y^t∣,WAPE=100∑t∣yt−y^t∣∑t∣yt∣\mathrm{MAE}=\frac{1}{H}\sum_{t=1}^{H}|y_t-\hat y_t|,\qquad \mathrm{WAPE}=100\frac{\sum_t|y_t-\hat y_t|}{\sum_t|y_t|}

Fit on the first 106 days and score only the final 14 days. No random train/test split. Compare against an unchanging last-week seasonal baseline.

Notation: t = days since January 1; pₜ = scheduled promotion flag; yₜ = units sold; β = fitted coefficients; H = holdout size. WAPE is undefined if all actuals are zero.

Python & C++ examples

Python dependencies: NumPy · scikit-learn
import numpy as np
from datetime import date, timedelta
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

# Synthetic daily units: same coffee fixture as the learning diagram.
day = np.arange(120)
promotion = (day % 19 < 3).astype(float)
units = np.floor(105 + 0.32 * day
    + 17 * np.sin(2 * np.pi * day / 7)
    + 7 * np.cos(2 * np.pi * day / 7)
    + 24 * promotion + 5 * np.sin(day * 2.37) + 0.5)

def features(t, promo):
    return np.column_stack([t / 100,
        np.sin(2 * np.pi * t / 7),
        np.cos(2 * np.pi * t / 7), promo])

X = features(day, promotion)
split = 106  # Final 14 days are unseen during fitting.
model = LinearRegression().fit(X[:split], units[:split])
pred = np.maximum(0, model.predict(X[split:]))
actual = units[split:]
# Promotions in the holdout are assumed scheduled and known at cutoff.
baseline = np.array([units[split - 7 + i % 7]
                     for i in range(len(actual))])
print("Synthetic holdout MAE:", mean_absolute_error(actual, pred))
print("Seasonal baseline MAE:", mean_absolute_error(actual, baseline))
print("Holdout WAPE (%):", 100 * np.abs(actual - pred).sum() / actual.sum())

# Refit only after evaluation. Future promotion is a planned scenario.
model.fit(X, units)
future = np.arange(120, 134)
planned_promotion = np.zeros(len(future))
forecast = np.maximum(0, model.predict(features(future, planned_promotion)))
for t, value in zip(future, forecast):
    print(date(2026, 1, 1) + timedelta(days=int(t)), round(float(value), 1))
# Real data: add rolling-origin backtests, stockout checks, uncertainty,
# and monitoring. This example makes no Google Cloud calls.

bigquery_forecast.sql
-- Replace project and dataset names. Running this incurs BigQuery costs.
-- Input: one daily row per product, with nonnegative units_sold.
CREATE OR REPLACE MODEL `your_project.sales.daily_forecast`
OPTIONS (
  MODEL_TYPE = 'ARIMA_PLUS',
  TIME_SERIES_TIMESTAMP_COL = 'order_date',
  TIME_SERIES_DATA_COL = 'units_sold',
  TIME_SERIES_ID_COL = 'product_id',
  DATA_FREQUENCY = 'DAILY',
  HOLIDAY_REGION = 'US'
) AS
SELECT order_date, product_id, units_sold
FROM `your_project.sales.daily_sales`
WHERE order_date < DATE '2026-04-17';

-- Evaluate these predictions against April 17–30 actuals first.
SELECT *
FROM ML.FORECAST(
  MODEL `your_project.sales.daily_forecast`,
  STRUCT(14 AS horizon, 0.9 AS confidence_level)
);

-- After evaluation, retrain on all available historical dates before
-- forecasting future dates. These are model-based intervals, not guarantees.

Production evaluation checklist

01Define the unit, daily grain, series ID, horizon, and decision the forecast supports.

02Audit returns, cancellations, missing dates, time zones, stockouts, and leakage from post-sale fields.

03Use rolling chronological backtests across several cutoffs, not a random split or one holdout alone.

04Compare against naive and seasonal-naive baselines; inspect error by product, store, and forecast horizon.

05Supply only future-known covariates. Scenario assumptions such as promotions must be explicit.

06Quantify uncertainty and choose metrics that match cost; WAPE can hide errors on low-volume products.

07Retrain on approved complete history after model selection, then monitor drift and realized error.

08Review privacy, access, regional availability, quotas, and current Google Cloud pricing before training.