Incomplete responses and observed-data estimation¶
Explicit data, explicit assumptions¶
Use IncompleteResponseMatrix for 0/1/NaN data. Zero is an observed
incorrect response; NaN is unobserved. ResponseMatrix keeps its complete
binary contract so that algorithms without a missing-data specification
cannot silently reinterpret NaN. Arrays are copied and read-only; item
labels may be a list or tuple. .aggregate() groups identical values and
masks, sums weights and removes zero-frequency rows. .to_complete()
rejects remaining NaN rather than dropping or imputing responses.
import numpy as np
from knowledgespaces import space_from_prerequisites
from knowledgespaces.estimation import (
BLIMConstraints, IncompleteResponseMatrix, estimate_blim_incomplete,
)
structure = space_from_prerequisites(["a", "b"], [("a", "b")])
data = IncompleteResponseMatrix(
["a", "b"], np.array([[0, 0], [1, 0], [1, 1], [1, np.nan]]),
counts=np.array([20, 40, 30, 10]),
)
fixed = BLIMConstraints(beta_fixed={"a": .1, "b": .1},
eta_fixed={"a": .1, "b": .1})
fit = estimate_blim_incomplete(structure, data, constraints=fixed)
print(fit.converged, fit.log_likelihood, fit.n_informative)
For observed item set \(O_r\), the likelihood contribution is
Missing response values are marginalized, not assigned zero or an estimated response. Dropping the missingness mechanism from the likelihood requires ignorability, including missing at random (MAR) and distinct response/ missingness parameters for direct likelihood inference. Rubin (1976) gives the conditions; they are assumptions, not properties established by passing a NaN array. MAR is relative to the variables in the model: ignoring auxiliary predictors of dropout does not become justified just because they exist elsewhere in the dataset.
de Chiusole et al. (2015) distinguish the MCAR IMBLIM from a model with state-dependent nonignorable missingness, MissBLIM. This API implements the observed-response likelihood. It does not fit MissBLIM, a general MNAR mechanism, or an automatic dropout correction. The label “missing-data support” must not be read as encompassing all these models. MD/MDML, SLM and IITA still require complete data in their current APIs.
EM updates and boundaries¶
The E-step conditions the latent state on observed cells. With expected state counts \(w_{rK}\), error updates use observed opportunities only:
Equality groups pool the numerators and denominators. Fixed parameters are
preserved, including zero; a fixed state prior may contain zero masses.
Free errors use the existing numerical box \([10^{-6},1-10^{-6}]\).
pi_init must be strictly positive for free priors so that a zero start
cannot lock a state out of EM. A probability-zero observed event raises an
error. The likelihood uses log probabilities, including exact fixed zeros.
A wholly missing row contributes probability one and log likelihood zero.
It is excluded from EM updates, so adding such rows does not change the
estimates or convergence speed. The result reports both n_respondents
(total weight) and n_informative (weight with any observed item). A sample
with no informative rows cannot identify a response model and is rejected.
An item never observed triggers UnobservedItemWarning; its unshared error
parameters retain their starts. A shared group can receive information
from other members. This warning is not a complete identifiability check.
objective_history records the initial and updated observed log likelihood.
Stopping uses maximum absolute parameter change; inspect converged and
n_iterations. A converged EM does not establish uniqueness or a global
optimum. npar counts free coordinates, not the dimension identifiable under
the observation design. AIC and BIC are conventional descriptive summaries;
BIC uses n_informative, and its regular-model interpretation is not
established by computing it. Fractional weights can be fitted, but are not
literal respondent counts for resampling.
Prediction, saturated likelihood and goodness of fit¶
fit.predict(patterns) and predict_blim_incomplete sum over missing
responses and return state posteriors. An entirely missing row retains the
prior. An impossible event has log probability \(-\infty\) and an undefined
(all-NaN) posterior. Rows with different masks can represent overlapping
events, so their probabilities are not disjoint multinomial cell masses.
Do not sum them or apply complete-table residuals to them.
from knowledgespaces.estimation import incomplete_gof
prediction = fit.predict(np.array([[1, np.nan], [np.nan, np.nan]]))
result = incomplete_gof(fit, data)
print(result.G2, result.saturated.converged)
saturated_incomplete fits a single free distribution over all \(2^{|Q|}\)
complete responses, observed through the supplied masks. It maximizes the
same marginal likelihood using multinomial EM, with uniform positive starts.
The objective is concave in that distribution, which need not be identified.
optimality_gap is \(\max_j (\partial\ell/\partial\pi_j)/N-1\), clamped
at zero. By concavity, \(N\) times this gap bounds the possible improvement
in log likelihood. Stopping checks that gap, not just small parameter updates.
This finite optimization check does not establish unique complete probabilities.
incomplete_gof returns twice the difference between saturated and model
observed log likelihoods, along with saturated convergence information.
A capped saturated fit can underestimate the difference; if its likelihood
is below the BLIM fit, G2 is returned as NaN with a warning. Comparison on
held-out observations is descriptive, not a training-data likelihood-ratio
test. No automatic chi-squared p-value or degrees of freedom is supplied:
the observable dimension depends on masks, constraints and regularity.
The saturated fit does not multiply by empirical mask frequencies. That
factorization would assume an independent mask distribution, stronger than
MAR. The separate independent_mask_pearson statistic uses precisely this
stronger assumption (or exogenous fixed mask strata), explicitly; it is not
a generic MAR goodness-of-fit test. Its normalization formula includes all
cells without constructing a full ternary response table. See the
bootstrap guide for the derivation and comparison boundaries.
max_patterns and max_memory_bytes guard explicit saturated enumeration;
BLIM fitting itself only needs observed patterns and specified states.
Memory preflights estimate arrays, not a hard process-memory cap.
Simulation and bootstrap¶
simulate_blim_incomplete(structure, observed_mask, ...) generates complete
individual BLIM responses, then retains cells where the boolean mask is True.
It preserves row order and uses label-aligned columns. This represents an
exogenous fixed observation design or MCAR, not generic MAR. Like the
complete simulator, it accepts error probabilities in \([0,1)\).
bootstrap_blim_incomplete requires an explicit mask_model:
"fixed"repeats each observed mask according to its integer frequency and simulates fresh responses independently of that mask. Its calibration requires an exogenous design/MCAR assumption."empirical_independent"draws masks from their weighted empirical law, independently of generated responses. Whole masks preserve within-row missingness dependence, while stratum sizes are resampled. The standalonesample_observation_masksfunction supplies this law for simulations.A callable receives read-only complete individual responses, item labels and the seeded random generator. It returns a boolean observed mask with the same shape. This permits a specified response-dependent mechanism; the software checks the contract, not its scientific appropriateness or ignorability. The mechanism must be compatible with the observed design.
Every replicate refits both BLIM and the observed saturated model using the same settings and constraints. Parameter and G2 arrays, iterations and convergence are retained. All-missing generated samples are recorded as failures with NaN parameter rows. An add-one Monte Carlo p-value is returned only if the observed fits and all replicates converge with finite G2. Otherwise p is NaN and failure/cap counts are reported; nothing is silently discarded. Parameter dispersion is descriptive and does not resolve nonidentifiability. A full missingness model is needed to justify inference under more general mechanisms; freezing masks does not solve that problem.
The description above is the default statistic="G2". With statistic="X2",
only fixed or empirical independent masks are accepted; arbitrary callbacks
are rejected. The statistic is the sum of Pearson statistics within the
exogenous mask strata. Saturated fitting is skipped: observed_gof is None,
G2 entries remain NaN and saturated iteration counts are zero. Calibration
requires finite Pearson scores and converged BLIM fits for all samples.
n_restarts and its initialization options apply the same random-start
policy to observed and simulated data. Every attempt remains in
estimate.restarts or restart_replicates; failed samples retain messages
in replicate_errors. See random starts. A finite raw tail
count remains available as descriptive_tail_fraction for inspecting capped
fits; it does not replace the unresolved calibrated p_value.
Probability individual data¶
See the dataset guide and cookbook/07_incomplete_probability.py.
The 159 wholly missing post rows provide no post-response information to a
post-only model. Its fit therefore equals the 345-completer fit; this does
not recover the post distribution among dropouts. Pre/post pairing is
preserved in the individual loader, enabling future explicitly specified
joint or conditional models instead of inventing pairings from frequencies.