A time series is stationary when its mean, variance, and autocorrelation structure stay stable over time, which is the condition most forecasting models assume. The practical workflow is straightforward: plot the series and its autocorrelation function, run an Augmented Dickey Fuller test alongside a KPSS test, and if either the visuals or the tests point to instability, apply a variance-stabilizing transform or differencing before modeling.
Key takeaways
| Point | Details |
|---|---|
| What it is | A stationary series has a stable mean, variance, and autocovariance structure that depends only on lag, not on calendar time. |
| Check visually first | Run-sequence, rolling mean/variance, and ACF plots catch problems a single test statistic can miss or misinterpret. |
| Run two tests, not one | ADF and KPSS test opposite null hypotheses; running both together is more informative than trusting either alone. |
| Match the fix to the problem | Unstable variance calls for a log or Box-Cox transform; a stochastic trend calls for differencing; over-differencing is a real risk. |
What Stationarity Really Means in Practice
Statisticians distinguish between two versions of this concept, and the difference matters more than it first appears. Strict stationarity requires that the entire joint probability distribution of the series remains unchanged under any shift in time, a condition so demanding that almost no real dataset satisfies it exactly. Weak stationarity, also called wide-sense stationarity, asks for much less: a constant mean, a constant variance, and an autocovariance between observations that depends only on the lag between them, never on the actual point in time. The NIST Engineering Statistics Handbook frames practical stationarity in exactly these terms, describing it as a series that looks stable in location and scale even while retaining meaningful autocorrelation.
For applied forecasting, weak stationarity is usually the target, and for good reason. Most classical models, including the autoregressive and moving-average structures behind ARIMA, are built on assumptions about constant variance and a lag-dependent covariance structure. When those assumptions hold even approximately, the autocorrelation function (ACF) settles into a predictable, decaying pattern that software and analysts alike can read and model. When they do not hold, the ACF often decays slowly or fails to settle at all, a signal worth returning to when running diagnostics.
A few recognizable patterns violate stationarity, and naming them helps you scan a plot with intent:
| Pattern | What it looks like | Typical example |
|---|---|---|
| Deterministic trend | Rises or falls along a fixed path | A company's revenue climbing steadily each quarter |
| Stochastic trend | A wandering, unpredictable drift; shocks accumulate rather than revert to a mean | Stock prices or exchange rates |
| Deterministic seasonality | Fixed, repeating cycles | Retail sales spiking every December |
| Volatility clustering | Periods of high variance followed by periods of calm, even when the mean looks stable | Financial returns |
Recognizing which of these patterns is present matters because the fix, and the modeling consequence, differs for each one.
Why Stationarity Shapes Forecasting and Model Choice
Stationarity is not a formality that precedes modeling — it is a structural requirement baked into the identification step of the Box-Jenkins method. The ACF and partial autocorrelation function (PACF) that analysts use to choose AR and MA orders are only interpretable when the underlying series is stationary, because those functions assume a fixed covariance structure. Feed a trending or seasonal series into that identification process, and the resulting ACF plot shows slow, uninformative decay instead of the sharp cutoffs that point toward a specific model order. The “I” in ARIMA exists specifically to solve this: differencing the series until it stabilizes, then modeling the stabilized version.
The distinction between deterministic and stochastic trends carries real consequences for forecast uncertainty. A series with a deterministic trend can be detrended by regression, and the residual uncertainty around that trend line stays roughly constant as the forecast horizon extends. A series with a stochastic trend, by contrast, accumulates variance the further out you forecast, because each new shock permanently shifts the level of the series. Treating a stochastic trend as if it were deterministic tends to produce forecast intervals that are too narrow and overconfident, especially several periods ahead.
Seasonality adds another branching decision. Analysts can remove it through seasonal differencing or model it explicitly with seasonal terms, and the right choice depends on whether the seasonal pattern is stable and worth preserving for interpretation. NIST’s guidance on Box-Jenkins model identification recommends regenerating the series and autocorrelation plots after each differencing step, checking that autocorrelation has meaningfully dropped and the mean has stabilized before moving forward.
A useful reference point: the NIST Handbook notes that weak stationarity concerns itself only with first and second moment behavior, meaning you are checking the mean, the variance, and the lag structure, not the full distribution. That narrower target is what makes the diagnostic workflow tractable for applied work.
Reading Visual Diagnostics Before You Trust a Test
Visual inspection comes first, not as a formality but because it catches problems that a single test statistic can miss or misrepresent. A run-sequence plot, simply the series values plotted against time, is the fastest way to spot an obvious trend, a sudden shift in level, or a variance that widens or narrows across the sample. NIST recommends this plot as a primary diagnostic precisely because it requires no statistical machinery and reveals patterns a p-value alone can obscure.
Four visual checks, used together, build a reliable early picture:
- Run-sequence plot: look for a rising or falling trend, an abrupt level shift, or a fan-shaped spread that signals changing variance.
- Rolling mean and rolling variance: compute these over a moving window (commonly 12 or 24 periods for monthly data) and check whether they drift or stay flat; a rolling mean that climbs steadily points to a trend, a rolling variance that grows points to heteroskedasticity.
- ACF plot: a stationary series typically shows autocorrelation that drops off quickly after the first few lags, while a slowly decaying ACF is a classic signature of a unit root or strong persistence.
- Seasonal subseries or spectral plots: these isolate periodic structure, such as a January-versus-July comparison across years, and help decide whether seasonal differencing or explicit seasonal modeling fits better.
The ACF deserves particular attention because it changes shape dramatically after a correct transformation. A series with a strong stochastic trend often shows an ACF that barely decays even at lag 20 or 30. After first differencing, if the transformation addressed the problem, the same ACF typically drops toward zero within the first few lags — a visible and often satisfying confirmation that the differencing worked. NIST’s model-identification guidance explicitly recommends regenerating these plots after each transformation step rather than assuming a fix worked based on theory alone.
Seasonal subseries plots serve a narrower but still important purpose: they separate a genuine periodic pattern, like a retailer’s predictable December surge, from random noise that merely looks cyclical over a short sample. When the periodicity is stable and strong, seasonal differencing (subtracting the value from the same period one cycle ago) usually resolves it faster than trying to model the seasonal shape directly. When the seasonal pattern is weak or shifts over time, explicit seasonal modeling terms, which preserve the pattern as an interpretable feature rather than discarding it, are often the better choice.
Statistical Tests: ADF, KPSS, and When They Disagree
Formal tests give you a number to report, but that number only means something if you know which hypothesis it is testing. The Augmented Dickey Fuller (ADF) test, documented in detail by statsmodels, tests the null hypothesis that a unit root is present, meaning the series is non-stationary. A small p-value here lets you reject that null and conclude the series is likely stationary. The test comes in three common regression specifications: no constant, a constant only, or a constant plus a deterministic trend, and the choice should match what you see in the run-sequence plot, since testing the wrong specification can produce a misleading result.
The Kwiatkowski-Phillips-Schmidt-Shin (KPSS) test, also documented by statsmodels, flips the logic entirely: its null hypothesis is that the series is stationary. A small p-value here means you reject stationarity — the opposite interpretation direction from ADF. KPSS offers two regression options, labeled “c” for level stationarity and “ct” for trend stationarity, and it relies on a Newey-West estimator to compute the long-run variance, a detail that affects results when lag truncation is chosen carelessly.
| Test | Null hypothesis | Small p-value means | Regression options |
|---|---|---|---|
| ADF (Augmented Dickey Fuller) | A unit root is present (non-stationary) | Reject the null; series is likely stationary | None, constant only, or constant plus trend |
| KPSS (Kwiatkowski-Phillips-Schmidt-Shin) | The series is stationary | Reject the null; series is likely non-stationary | "c" for level stationarity, "ct" for trend stationarity |
Running both tests together, rather than leaning on just one, is the practical rule worth following:
- ADF rejects, KPSS fails to reject: both tests agree the series is stationary, the strongest possible signal.
- ADF fails to reject, KPSS rejects: both agree the series is non-stationary, pointing toward differencing or detrending.
- The two tests disagree: this happens often enough that it should not be treated as a software error; it usually signals borderline persistence, a structural break, or a specification mismatch between the two tests.
ADF and KPSS test opposite null hypotheses, which means a given p-value supports entirely different conclusions depending on which test produced it — a point statsmodels’ documentation makes explicit by requiring users to state the regression type for each test before interpreting results.
Phillips-Perron is a useful complement to ADF because it adjusts for serial correlation and heteroskedasticity non-parametrically rather than through added lag terms, which makes it more robust when the error structure is irregular. When a structural break, such as a policy change or a one-time shock, might be distorting a standard unit-root test, the Zivot-Andrews test accounts for a single unknown breakpoint and often reverses conclusions that looked clear-cut under ADF alone. Analysts who need to compare the two should check the p-value explainer for a refresher on what these thresholds do and do not tell you, since a p-value near 0.05 under either test deserves more scrutiny than a single rule of thumb can provide.
The practical discipline that ties all of this together: always report which null hypothesis you tested, run at least two complementary tests rather than one, and never let a borderline p-value override what the run-sequence and ACF plots already showed you.
Transformations That Remove Non-Stationary Behavior
Different violations of stationarity call for different fixes, and applying the wrong one wastes data or leaves the real problem untouched. Forecasting: Principles and Practice treats differencing as the standard first tool for a stochastic trend, computing the first difference — the current value minus the previous value — the difference between each observation and the one before it. First differencing removes a linear trend effectively but costs one observation at the start of the series, a small price for short series but worth tracking.
Seasonal differencing works the same way but subtracts the value from one full cycle earlier: the current value minus the value from m periods ago, where m is the seasonal period (12 for monthly data with yearly seasonality, 4 for quarterly). fpp3’s chapter on stationarity notes that some series need both seasonal and first differencing applied in sequence, since removing seasonality alone can still leave an underlying trend in place.
A different family of problems, unstable variance rather than unstable mean, calls for a different tool entirely:
- Log transform: compresses large values more than small ones, a common fix when variance grows alongside the level of the series, such as in sales or population data.
- Square-root transform: a gentler version of the log transform, useful when variance grows but not as sharply.
- Box-Cox transform: a flexible family that includes log and square-root as special cases, letting the data itself suggest the best stabilizing power.
- Handling nonpositive values: all three transforms require positive data, so a small constant is often added before transforming when the series contains zeros or negative values.
Detrending by regression is worth distinguishing from differencing because it answers a different question. Fitting a line or curve to the series and subtracting it removes a deterministic trend while preserving the original scale and interpretability of the residual, which matters when you want to talk about deviations from an expected path rather than period-to-period changes. Differencing, by contrast, is the right tool when the trend is stochastic, since detrending a random walk leaves spurious structure behind.
A sensible default order, though not a rigid rule, is to address variance first, then trend, then seasonality: stabilize the variance with a log or Box-Cox transform, difference to remove trend, then apply seasonal differencing or seasonal terms if periodicity remains. NIST’s Box-Jenkins guidance frames this as an iterative checklist rather than a one-pass recipe, because the effect of each transformation should be checked with fresh plots before deciding whether another step is needed.
A Step-by-Step Workflow You Can Apply to Any Series
The diagnostic process described above comes together into a repeatable sequence. Following these steps in order, and retesting after each change, keeps the process from turning into guesswork.
- Plot the raw series and compute rolling mean and variance over a sensible window, looking for trend, level shifts, or widening spread; the exploratory data analysis workflow covers the plotting habits this step relies on.
- Run ADF and KPSS together, specifying the regression type that matches what the plot showed (constant only, or constant plus trend), and note which null each test rejected or failed to reject.
- If variance looks unstable, apply a log or Box-Cox transform first, then rerun both tests before touching the trend, since stabilizing variance can change how the trend appears.
- Check for seasonality using subseries or ACF plots at seasonal lags; if present, apply seasonal differencing or add explicit seasonal terms depending on whether the pattern is worth preserving.
- Apply first differencing only if unit-root evidence persists after the previous steps, then regenerate the run-sequence and ACF plots to confirm the autocorrelation has dropped off.
Stop once the ACF and PACF settle into a stable, interpretable pattern and the rolling statistics look flat, rather than continuing to difference in search of a perfect test result. Over-differencing is a recognized pitfall precisely because it introduces artificial negative autocorrelation and removes signal that the model needed, a caution fpp3 raises directly when discussing how many differences a series actually needs.
Pitfalls That Quietly Undermine a Stationarity Check
Pitfalls to check for before you trust a result
- Over-differencing Applying a second or third difference when the first already resolved the trend adds noise and makes the series harder, not easier, to model.
- Structural breaks A single sharp shift in level or trend can make ADF and KPSS both point toward non-stationarity even when the series is stable within each regime; a breakpoint-aware test such as Zivot-Andrews often gives a clearer answer.
- Variance clustering Periods of high and low volatility, common in financial returns, violate the constant-variance assumption even when the mean is flat, and call for ARCH or GARCH modeling rather than further differencing.
- Small samples and lag choices With limited data, the lag length chosen for ADF or the truncation used in KPSS's long-run variance estimate can shift the p-value meaningfully, so short series deserve a hedged interpretation.
Continuing the Diagnostic Work Beyond This Guide
Learn and Calculators exist specifically so you do not have to choose between understanding a concept and actually applying it. The plain-English explanations on stationarity connect directly to interactive calculators for rolling statistics and p-value interpretation, which means you can test a real dataset without writing the underlying formulas yourself. For readers working through a class assignment or a research project, document every transformation and test specification as you go, since reproducibility matters as much in applied statistics as the correctness of any single test. The Applied Statistics hub carries worked examples that extend this workflow into full forecasting problems.
When Is Approximate Stationarity Good Enough?
Chasing a perfectly stationary series is often the wrong goal. Strict stationarity rarely exists in real data, and insisting on it tends to push analysts toward over-differencing, which trades away interpretable structure for a cleaner-looking test result. The better question is whether the series is stationary enough for the model and the forecast horizon you actually care about.
Mild heteroskedasticity, for instance, matters far more for a 24-month financial forecast than for a short-horizon operational one, where the variance is unlikely to shift enough to distort the result. Seasonal structure, similarly, is sometimes worth preserving as an interpretable feature rather than differencing away, especially when stakeholders need to see the seasonal pattern directly rather than infer it from a transformed series. What matters for any analyst, student or professional, is transparency: report which tests you ran, which null each one addressed, and which transformations you applied in what order, so another reader can judge whether your tolerance for imperfection was reasonable.
Where to Test and Transform Your Own Time Series
Reading about ADF and KPSS tests is one thing; running them against a dataset you actually care about is another. Each statistics guide moves from the definition of a rolling variance to computing one on your own numbers without switching tools or learning a new programming language.
For a dataset with more structure, like a seasonal sales series or an economic indicator, the Applied Statistics hub walks through full worked examples that apply the diagnostic workflow described here start to finish. If you are newer to the field and want the conceptual groundwork before diving into tests and transforms, the fundamental statistics guide is a reasonable place to start, and from there the calculators landing page gives you the tools to put any of it into practice on your own data.
Recommended
- Exploratory Data Analysis: A Practical Workflow
- Linear Regression Assumptions: What They Are & How to Check
- Data Drift Detection: A Practical Guide for Production ML Systems
Sources
Sources
- Stationarity — NIST/SEMATECH e-Handbook of Statistical Methods NIST
- Forecasting: Principles and Practice (3rd ed.) — Stationarity and Differencing Hyndman & Athanasopoulos
- statsmodels.tsa.stattools.adfuller statsmodels
- statsmodels.tsa.stattools.kpss statsmodels
- Box-Jenkins Model Identification — NIST/SEMATECH e-Handbook NIST
- Box-Cox Normality Plot — NIST/SEMATECH e-Handbook NIST
FAQ
Frequently asked questions
- What are the four main components of a time series?
- A time series is typically broken into trend (the long-term direction), seasonality (regular, fixed-period cycles), cyclical variation (longer, irregular fluctuations tied to broader conditions), and the remainder or irregular component (random noise left after removing the others). Knowing which components are present helps decide whether to difference, detrend, or model seasonality explicitly.
- What is the Dickey Fuller test?
- The Dickey Fuller test, and its more common extension the Augmented Dickey Fuller test, checks whether a time series has a unit root, meaning it is non-stationary. Its null hypothesis is that a unit root is present, so a small p-value lets you reject that null and conclude the series is likely stationary.
- What is a wide-sense stationary process?
- A wide-sense, or weakly, stationary process is one where the mean and variance stay constant over time and the covariance between any two points depends only on the distance, or lag, between them, not on when they occur. The NIST Handbook treats this as the practical standard for most applied time series work, since it is far easier to verify than full strict stationarity.
- How do you test for stationarity?
- Start with a run-sequence plot and rolling mean and variance to spot obvious trend or changing spread, then run the ADF and KPSS tests together, since they test opposite null hypotheses and give a more complete picture than either alone. If the two tests disagree or the series shows a sudden shift in level, check for a structural break with a test like Zivot-Andrews before deciding on a transformation.
- Should I use ADF or KPSS if they give conflicting results?
- Conflicting results between ADF and KPSS are common and usually point to borderline persistence, a structural break, or a mismatched regression specification rather than an error. Go back to the run-sequence and ACF plots, check whether a break in the series might explain the disagreement, and consider a breakpoint-aware test rather than trusting either p-value in isolation.