Skip to contents

Creates an advanced, publication-ready two-panel dashboard for visualizing predicted values and highlighting the most notable cases or strata. What "notable" means depends on the model type, and the labelled points are not statistical outliers in the regression-diagnostic sense:

  • Gaussian and Poisson (and the ordinal "expected_score" mode): the cases/strata whose prediction sits furthest from the mean prediction (largest deviation), ranked by absolute deviation.

  • Binomial: the cases/strata with the largest absolute deviance residual, i.e. where the observed 0/1 outcome – for an aggregated binomial (cbind(successes, failures), a proportion with its trial counts as prior weights, or a brms y | trials(n) fit), the observed successes out of trials – is least consistent with the fitted probability (worst-fit points), ranked by \(|deviance residual|\). As in R's binomial family, a prior weight counts as that many trials, so a 0/1 outcome with weights is scored as that many successes or failures. Every row's outcome is coded as the model coded it – by label, whatever the order of a factor's levels in data – and a factor with more than two levels, as in R's binomial family, as its first level against all the others. Predictions are per-trial probabilities for every binomial fit. A row whose outcome or trial count is unknown has no residual: it is left out of the stratum residual means, is never labelled, and is drawn as a hollow point.

  • Ordinal "surprise" mode: the cases/strata with the highest surprise \(-\log P(\text{observed category})\), i.e. the least probable observations under the model. It needs the observed response in data: a row whose response is missing is left out, and one whose category is not among the model's fitted categories is left out with a warning.

Usage

plot_prediction_deviation_panels(
  model,
  data = NULL,
  type = c("auto", "gaussian", "poisson", "binomial", "ordinal"),
  ordinal_mode = c("surprise", "expected_score"),
  top_n_labels = 5,
  strata_info = NULL
)

Arguments

model

A fitted model object (e.g., from `lm()`, `glm()`, `MASS::polr()`, or `lme4::glmer()`).

data

The original data frame used to fit the model. If `NULL`, attempts to extract from the model. When supplied, everything is computed for its own rows: their predictions, their binomial residuals (from their own outcomes and trial counts) and their weights in the stratum summaries. A row's weight is read from its column in `data` – the sampling-weight column, or the column a bare model's `weights =` argument names (`glm()`, `lmer()`, `MASS::polr()`, ...) – or else from the fitted row with the same row name and prediction (a `fit_maihda()` model stores its weights as values). If any plotted row's weight cannot be found, every row is weighted equally, with a warning.

type

Model type: "auto" (default), "gaussian", "poisson", "binomial", or "ordinal".

ordinal_mode

For ordinal models: "surprise" (default, based on observation probability) or "expected_score".

top_n_labels

Number of points to label on the plot. The ranking metric depends on the model type (see Description): deviation from the mean prediction for Gaussian/Poisson and the ordinal expected-score mode, absolute deviance residual for binomial, and surprise for the ordinal surprise mode. Default is 5.

strata_info

Optional data frame of strata labels, generally extracted from `maihda_model` objects.

Value

A `patchwork` object containing two `ggplot2` panels.