Skip to contents

Makes predictions from a fitted MAIHDA model, either at the stratum level or individual level.

Usage

predict_maihda(
  object,
  newdata = NULL,
  type = c("individual", "strata", "response", "link"),
  scale = c("response", "link"),
  allow_new_levels = FALSE,
  ...
)

Arguments

object

A maihda_model object from fit_maihda().

newdata

Optional data frame for making predictions. If NULL, uses the original data from model fitting.

type

Character string specifying prediction type:

  • "individual": Individual-level predictions including random effects

  • "strata": Stratum-level predictions (random effects only). For a longitudinal (growth-curve) fit a stratum is a trajectory, so this returns the per-stratum trajectory parameters (baseline deviation, random intercept and random slope(s)) rather than a single random effect. A non-longitudinal fit whose stratum random effects include random slopes (a hand-written (1 + x | stratum)) is an error here, as in summary(): the single value this would return is the intercept alone – the stratum effect only where the slope variables are zero.

For backward compatibility, "link" or "response" may also be passed here and will be interpreted as individual-level predictions on that scale.

scale

Character string specifying the prediction scale for individual-level predictions: "response" (default) or "link". For a cumulative (ordinal) model the "link" scale is the latent location \(\eta\) and the "response" scale is the expected category score \(\sum_k k P(Y = k)\) (categories scored 1..K in their declared order). For an aggregated-binomial fit (an lme4 cbind(success, failure) or a brms success | trials(n)) the "response" scale is the per-trial probability on both engines (not the expected success count).

allow_new_levels

Logical. By default (FALSE) a stratum in newdata that the model never saw – whether supplied directly as a stratum column or rebuilt from the grouping variables – is an error, for every engine, matching lme4's default. Set TRUE to instead predict unseen strata with the stratum random effect dropped (treated as zero), while keeping any other random effect the row participates in (e.g. a contextual (1 | school) intercept from fit_maihda(context = ), or a longitudinal growth term) – the same behaviour as lme4's allow.new.levels, which zeroes only the unseen level's effect and keeps seen ones. For the usual single-stratum model the stratum is the only random effect, so the result is the fixed-effects-only prediction, evaluated at a zero random effect. That is a conditional (stratum-specific) prediction for a stratum whose effect happens to be zero; it is not a response-scale population average (marginal mean), which requires integrating over the random-effect distribution. The two coincide on the link scale, and on the response scale only under the Gaussian identity link. Under a log link with stratum variance \(\tau^2\) the marginal mean is larger by a factor \(\exp(\tau^2/2)\), and under a logit link the marginal probability is attenuated towards 0.5 (Nakagawa, Johnson & Schielzeth 2017). Because the inverse link is monotone and the random effect symmetric about zero, the response-scale value returned here is the median of the stratum-specific means across the random-effect distribution, not their average. This affects type = "individual" only: a stratum-level prediction (type = "strata") has no random effect to report for an unseen stratum, so unseen strata remain an error there regardless.

...

Additional arguments passed to predict method of underlying model.

Value

Depending on type:

  • For "individual": A numeric vector of predicted values on the requested scale, one per row of newdata – or, when newdata is NULL, one per ANALYTIC row, matching nrow(object$data). A model fitted with na.action = na.exclude is not padded back out to the original input rows the way predict() on the underlying fit is, so the result can always be bound to object$data.

  • For "strata": A data frame with stratum ID and predicted random effect. When newdata is supplied, the result is restricted to the strata present in newdata (and a stratum the model never saw is an error, as for "individual"); when newdata is NULL, every training stratum is returned.

Examples

# \donttest{
strata_result <- make_strata(maihda_sim_data, vars = c("gender", "race"))
model <- fit_maihda(health_outcome ~ age + (1 | stratum), data = strata_result$data)

# Individual predictions
pred_ind <- predict_maihda(model, type = "individual")

# Stratum predictions
pred_strata <- predict_maihda(model, type = "strata")
# }