Statohub Browse calculators
Machine Learning Statistics Practitioner guide

Interpret Probabilities, Not Odds: Logistic Regression

This logistic regression interpretation guide turns coefficients into odds ratios, marginal effects, and predicted probabilities, verified with diagnostics.

By Statohub Editorial Team Published October 2026Reviewed October 202612 min read

Logistic regression coefficients represent changes in log-odds, and exponentiating a coefficient converts it into an odds ratio, the multiplicative change in the odds of the outcome per one-unit increase in a predictor. Because odds are not probabilities, we recommend pairing every coefficient with its exponentiated confidence interval and, when the audience needs real-world meaning, a marginal effect or predicted probability. Model fit and diagnostics determine whether any of these numbers deserve trust.

Key takeaways

Point Details
Odds ratios are multiplicative, not constant Each one-unit increase multiplies the odds by exp(β), but the same ratio produces different probability swings depending on the baseline probability.
Categorical coefficients need a reference level Each coefficient compares one category level to its baseline, so exp(β) only means something once you state what it is being compared to.
The intercept is a log-odds baseline β0 is the log-odds when every predictor sits at zero or its reference level; centering predictors turns it into the log-odds for an average-profile case.
A confidence interval spanning 1 means no clear effect Odds-ratio confidence intervals that include 1 are compatible with no association, and odds ratios only approximate relative risk when the outcome is rare.
Diagnostics decide whether to trust the coefficients Check linearity in the logit, separation, multicollinearity, and calibration before interpreting any coefficient as real.

What the model estimates: log-odds, odds, and probability

A logistic regression model does not predict probability directly. It predicts the logit, the natural log of the odds of the event, as a linear function of the predictors: logit(p) = ln(p / (1 minus p)) = β0 + β1x1 + … + βpxp. This structure, formally the logit model, keeps predicted probabilities bounded between 0 and 1, something ordinary linear regression cannot guarantee for a binary outcome, as the CASRAI guide to logistic regression explains.

To move from the linear predictor back to something intuitive, we exponentiate it to get odds, then apply the inverse logit to recover probability: p = 1 / (1 + exp(negative linear predictor)). Because this transformation is a curve rather than a straight line, the same coefficient produces a larger probability swing near p = 0.5 than it does near p = 0.05 or p = 0.95. That nonlinearity is the reason raw coefficients can mislead readers who expect a constant effect.

From linear predictor to probability Four-step flow: the linear predictor is exponentiated into odds, converted via the inverse logit into a bounded probability, with nonlinearity producing a larger probability change near 0.5 than near the extremes. 1 Linear predictor logit = linear function 2 Odds exp(linear predictor) 3 Probability inverse-logit, 0-1bounded 4 Nonlinearity larger change near 0.5
Figure 1. The logit transformation moves from a linear predictor through odds to a bounded probability — nonlinear throughout.

Reading coefficients: continuous vs. categorical predictors

For a continuous predictor, the coefficient β is the expected change in log-odds for a one-unit increase, holding other variables constant; exp(β) converts that into an odds multiplier for the same one-unit increase. When a unit is awkward (age in days, income in dollars), we suggest reporting the effect per meaningful increment instead, exponentiating β multiplied by that increment (say, per 10 units) rather than the raw per-unit value, a practice outlined in the Illinois guide to slope and intercept interpretation.

Categorical predictors work differently. Each coefficient reflects the log-odds difference between one level and an explicit reference category, so exp(β) is the odds ratio comparing that level to the reference rather than a standalone effect, as Minitab’s documentation on binary logistic coefficients notes. Always state the reference level in your write-up, and document how each factor was coded before anyone tries to sign-check your results.

Intercept and reference levels: what β0 really means

The intercept β0 is the log-odds of the event when every continuous predictor equals zero and every categorical predictor sits at its reference level, and exp(β0) gives that baseline odds, according to the Illinois exploration notes on logistic regression. When zero is not a realistic value, like age zero or income zero, the intercept becomes a mathematical artifact rather than a meaningful baseline. Centering continuous predictors (subtracting their mean before fitting) turns the intercept into the log-odds for an average-profile individual, which is usually far more useful to report.

Odds ratios, confidence intervals, and p-values

Once you exponentiate a coefficient to get an odds ratio, exponentiate the confidence interval endpoints too, producing an interval on the odds-ratio scale rather than the log-odds scale. An interval that spans 1 is compatible with no multiplicative association between that predictor and the odds of the outcome, a point emphasized in the UCLA FAQ on interpreting odds ratios.

A few practical rules keep this honest:

  • A p-value tests the coefficient against a null hypothesis conditional on the specified model; it says nothing about effect size or predictive accuracy.
  • Odds ratios approximate relative risk only when the outcome is rare (roughly below 10%); for common outcomes, predicted probabilities or risk differences communicate the effect more faithfully.
  • Outcome coding changes everything: reversing which category counts as the event flips coefficient signs and turns each odds ratio into its reciprocal, so confirm the event coding before interpreting any sign.

Marginal effects and predicted probabilities: translating to the probability scale

Odds ratios are multiplicative and constant across the predictor’s range, but the probability change they produce is not. Marginal effects solve this by reporting the expected change in probability for a one-unit change in a predictor, computed in several ways: the average marginal effect (AME) across all observations, the effect at the sample mean, or the effect at chosen representative values, as the University of Colorado Denver course notes on marginal effects describe.

Two predictors with identical odds ratios can produce very different absolute probability changes depending on where the baseline sits, which is exactly the nonlinearity described earlier. We recommend reporting predicted probabilities for a handful of representative profiles alongside any odds ratio, especially once interactions enter the model.

Marginal effects: key points Four ways to summarize a marginal effect: the average marginal effect across observations, the effect at the sample mean, the effect at chosen representative values, and reporting predicted probabilities for example profiles. Marginal effect types AME Average effect across observations At-mean Effect at the sample mean At-values Effect at chosen representative values Predicted probs Report probabilities for example profiles
Figure 2. Four common ways to compute and report a marginal effect, each answering a slightly different question.

Interactions and conditional effects

Adding an interaction term changes what a main-effect coefficient means. Once a predictor interacts with another, its main-effect coefficient only describes the effect when the interacting variable equals zero, not an effect that holds everywhere, as the margins literature cautions. Interpreting an interacted model correctly means computing the full linear predictor for the subgroup of interest first (main effects plus the relevant interaction terms), then exponentiating or converting to probability, rather than exponentiating each coefficient in isolation.

For communicating interaction results, margins and plotted predicted probabilities across the range of one variable, split by levels of the other, tend to convey the conditional pattern far more clearly than a table of individually exponentiated coefficients.

Diagnostics and model adequacy you must report

A coefficient’s interpretation is only as trustworthy as the model producing it. Several checks belong in any serious write-up, drawn from standard logistic regression diagnostic guidance — including a screen for multicollinearity using variance inflation factors and a check of the standard regression assumptions that carry over to the logit scale:

Diagnostics to run before trusting a coefficient

  • Check linearity in the logit Use component-plus-residual plots or splines for continuous predictors instead of assuming a straight-line relationship.
  • Watch for separation Small or empty cells, or a predictor that perfectly predicts the outcome, produce unstable or infinite estimates — consider a penalized likelihood method such as Firth's correction.
  • Screen for multicollinearity and influence Check variance inflation factors, examine influential observations through leverage and Cook's distance, and verify standard regression assumptions on the logit scale.
  • Assess calibration and discrimination separately Compare predicted probability to observed frequency (calibration) and check ROC/AUC (discrimination) — a model can do well on one and poorly on the other.

Statistical significance on its own never confirms that a predictor matters practically; a well-estimated but tiny effect can be significant in a large sample and irrelevant in practice.

Short worked example: read an output and state three clear interpretation statements

Say a model predicts loan default from income (in thousands) and a binary variable for prior default history (reference: no prior default). Suppose the output gives a coefficient of negative 0.04 for income and 1.10 for prior default history, with an intercept of negative 1.2.

  1. For income, exp(negative 0.04) is about 0.96, so each additional $1,000 of income is associated with roughly a 4% reduction in the odds of default, holding prior history constant.
  2. For prior default history, exp(1.10) is about 3.0, so someone with a prior default has about three times the odds of defaulting again compared to someone without one, at the same income.
  3. For a borrower with $50,000 income and no prior default, the linear predictor is negative 1.2 plus (negative 0.04 times 50) equals negative 3.2, giving a predicted probability of about 1 / (1 + exp(3.2)), roughly 0.04, or a 4% chance of default.
Reading the worked example's coefficients as odds ratios
Predictor Coefficient (β) exp(β) Interpretation
Income (per $1,000) -0.04 0.96 ~4% lower odds of default per additional $1,000 of income
Prior default history (vs. none) 1.10 3.0 ~3x the odds of default vs. no prior default
Intercept -1.2 n/a Baseline log-odds at $0 income, no prior default

A quick diagnostics check, confirming no separation, acceptable VIF values, and reasonable calibration, should accompany any of these three statements before they go into a report.

Put together, these three statements trace the full interpretation chain: a raw coefficient becomes an odds ratio through exponentiation, an odds ratio becomes a probability through the inverse logit, and neither number means much on its own until the model behind it has passed the diagnostics checklist above.

How Statohub helps: calculators and applied guides to check your interpretation

Our Applied Statistics hub walks through examples like this one in full, and our calculators let you verify each step rather than trust arithmetic alone. We built our guides and tools to sit side by side, so a concept explained in a Learn article connects directly to a calculator built for that exact computation.

Reader perspective: common interpretation mistakes students make

The mistake we see most often is treating an odds ratio as if it were a risk ratio, which overstates the effect whenever the outcome is not rare. A close second is quoting a coefficient without naming the reference level, which leaves the reader guessing what “higher” even means. The third is reporting a significant coefficient from a model nobody checked for separation or calibration. Running the marginal effects and glancing at a predicted-probability table for a few representative cases catches most of these before they reach a final report.

Reading an output table is one skill; reproducing the numbers yourself is another, and that second step is where interpretation actually sticks. Our Calculators page includes a probability calculator for converting log-odds to probability by hand, and the Applied Statistics section carries worked examples that extend the one above into full datasets. The Learn, Calculate, Apply sequence on Statohub exists specifically so a formula you read about becomes a number you can check yourself, then a decision you can defend.

Sources

Sources

  1. Logistic regression (logit model) CASRAI
  2. Interpreting the estimated coefficients in binary logistic regression Minitab Support
  3. Slope and intercept interpretations — Exploration: Statistical Learning University of Illinois
  4. How do I interpret odds ratios in logistic regression? UCLA Statistical Consulting
  5. Logistic Regression — Stata Data Analysis Examples UCLA Statistical Consulting
  6. What is complete or quasi-complete separation in logistic/probit regression, and how do we deal with them? UCLA Statistical Consulting
  7. "P < 0.05" Might Not Mean What You Think: American Statistical Association Clarifies P Values PMC / American Statistical Association

FAQ

Frequently asked questions

Can you explain logistic regression in a simple way?
Logistic regression estimates the probability of a binary outcome (yes or no, event or no event) by modeling the log-odds as a straight-line combination of predictors, then converting that back to a probability between 0 and 1. Each predictor's coefficient tells you how the log-odds shift, and exponentiating it tells you how the odds multiply.
What does a P-value of 0.05 mean in regression analysis?
A p-value of 0.05 means that, assuming the coefficient is truly zero and the model is correctly specified, a result this extreme or more extreme would occur about 5% of the time by chance. It does not measure how large or important the effect is, only how compatible the data are with no effect.
How do you interpret an odds ratio?
An odds ratio above 1 means the odds of the event increase as the predictor increases (or for that category versus the reference); below 1 means the odds decrease. Confirm which outcome is coded as the event first, since reversing that coding flips the ratio to its reciprocal.
What is the difference between logistic and linear regression?
Linear regression predicts a continuous outcome directly as a straight-line function of predictors, while logistic regression predicts the log-odds of a binary outcome and converts that back to a bounded probability. A linear regression coefficient is a direct change in the outcome, while a logistic coefficient is a change in log-odds that must be exponentiated or converted to probability to interpret intuitively, as covered in our guide to regression assumptions.