Prediction and reliability

BLIMEstimate.predict computes response probabilities and state posteriors from a fitted model. BLIMEstimate.reliability computes population measures for administration of every item, keeping the state and parameters fixed.

from knowledgespaces.datasets import load_probability
from knowledgespaces.estimation import estimate_blim

ds = load_probability(wave="pre", model="K1")
fit = estimate_blim(ds.structure, ds.data, method="MD")
prediction = fit.predict(ds.data.patterns, items=ds.data.items)
print(prediction.probabilities)
print(prediction.posteriors.shape)  # patterns by states

reliability = fit.reliability()
print(reliability.ks_reliability, reliability.rp_reliability)
print(reliability.accuracy, reliability.adjusted_accuracy)
print(reliability.expected_discrepancy)

The functions estimation.predict_blim and metrics.blim_reliability also accept supplied parameters, without fitting. Item mappings and state mappings avoid ambiguity about ordering. Array parameters follow the sorted domain and canonical state order; prediction’s explicit items sets its response-column and item-parameter order. Fitted methods align the fit’s parameters by label automatically. Returned arrays are read-only snapshots with accompanying labels.

Entropy measures (2024)

Let \(K\) be the true state, \(R\) the full response pattern, and \(p(K,R)=\pi_K P(R\mid K)\) their joint distribution. The implementation evaluates Equations 10–11 of de Chiusole et al. (2024):

\[\mathrm{RP}=\frac{I(K;R)}{H(R)},\qquad \mathrm{KS}=\frac{I(K;R)}{H(K)}.\]

All entropies use bits. The result includes \(H(R)\), \(H(K)\), \(H(R\mid K)\) and \(I(K;R)=H(R)-H(R\mid K)\). These are model-based quantities. Plugging in fitted parameters provides point estimates and does not account for uncertainty about parameters or the structure.

Boundaries, cost and interpretation

The probability functions accept closed-box error probabilities in \([0,1]\) and zero state probabilities to evaluate deterministic limits. This extends the calculation to boundaries beyond the positive-prior, informative-item setting used for the 2025 interpretation. A zero entropy denominator produces NaN, as does adjusted accuracy when max(pi)=1. Conditional state accuracy is NaN for a state of prior probability zero. An impossible response has probability zero, log probability -inf, and undefined (NaN) posterior; the library does not invent a posterior.

Reliability sums over all \(2^{|Q|}\) response patterns in chunks. The default max_patterns=1_048_576 limits enumeration work; a conservative allocation estimate is checked against max_memory_bytes. Lowering chunk_size reduces working memory, while enumeration time stays exponential. The finite sums use floating-point arithmetic, not symbolic exactness.

These indices do not establish model fit, predictive performance on new data, or the reliability of an adaptive stopping rule. Prediction at supplied patterns does not enumerate unobserved patterns. See scope and assumptions for the distinction between theoretical assumptions and the current implementation’s coverage.