Statistical reports, discrepancy and expected errors

blim_report brings together fitted coefficients and constraints, marginal mastery, expected errors, stored goodness of fit, optional minimum discrepancies and optional local BLIM trade-offs. It returns an immutable result and a detached JSON-ready export. It does not refit the model or construct new inference from these descriptive quantities.

import json

from knowledgespaces.datasets import load_doignon_falmagne_7
from knowledgespaces.estimation import blim_report, estimate_blim

dataset = load_doignon_falmagne_7()
fit = estimate_blim(dataset.structure, dataset.data, method="MD")
report = blim_report(fit, dataset.data)
assert report.discrepancy.training_data is True
assert report.aicc_sample_size == 1000
assert report.mastery.shape == (5,)
export = report.to_dict()
serialized = json.dumps(export, allow_nan=False, indent=2)
# To save explicitly: Path("report.json").write_text(serialized)

The item order is report.items; report.states records the state labels. report.coefficients contains one row per merged beta/eta equality group, with its fixed/free status, and one row per actual state probability. State probabilities are not a list of independent simplex coordinates. For an SLM report, solvability parameters g are included and state probabilities are marked derived=True. report.n_parameters is the fitted model’s nominal free-parameter count, not an identified dimension inferred by this report.

Marginal mastery and full-test errors

For each item q the report computes

[ m_q=\sum_{K:q\in K}\pi_K,\qquad E[S_q]=\beta_q m_q,\qquad E[G_q]=\eta_q(1-m_q). ]

mastery gives m, expected_slips gives E[S], and expected_guesses gives E[G]. Summing the latter two arrays gives expected_total_errors: the expected number of errors for a new respondent completing the entire test under the fitted population distribution. This reproduces the quantities reported as nerror by pks, with per-item contributions retained.

These quantities do not condition on a particular response pattern. They also do not multiply by the training sample size. For posterior state probabilities conditional on a supplied pattern, use fit.predict(...). For an incomplete-data fit the report’s error expectations still describe a hypothetical complete test; they are not expected errors among observed cells under a missingness design.

Minimum discrepancy distribution

from knowledgespaces.estimation import minimum_discrepancy

distance = minimum_discrepancy(fit, dataset.data, chunk_size=8)
assert distance.counts.sum() + distance.infeasible_weight == distance.total_weight
assert len(distance.distances) == dataset.data.n_patterns
assert len(distance.counts) == dataset.data.n_items + 1

For each complete response row R, the distance is the minimum of |R △ K| over admissible model states. counts[d] is the supplied frequency or weight at distance d; the vector includes zero bins through the number of items. distances retains input row order, and mean is frequency-weighted.

As in pks getMD, a declared fixed-zero slip or guess forbids the corresponding error event before minimizing. Equality groups propagate the fixed-zero restriction. A parameter estimated numerically as zero, or a zero prior mass on a state, does not by itself change this distance rule. The excess inclusion radius and hyperbolic MD weights affect estimation assignments, not the minimum Hamming distance being reported here.

If every state is forbidden for a response, its distance is infinity. infeasible_weight retains its weight; a positive infeasible weight makes the mean infinite. An infeasible zero-frequency row does not make the weighted mean infinite. The export represents nonfinite numbers as JSON null while retaining the infeasible weight explicitly.

The input must be a complete ResponseMatrix, and its mutable arrays are revalidated. No missing cell is interpreted as an incorrect response. Chunking bounds working allocations without enumerating all 2^|Q| response patterns; max_memory_bytes is an allocation guard, not a process-RSS limit.

Training quantities and other data

The fit’s likelihood, GOF, AIC and BIC are always under export["training"]. Supplying different data to the report computes their discrepancy only; it never relabels the stored training GOF as a test on those data. discrepancy.training_data is True when the supplied labeled empirical measure matches the fit’s training fingerprint, False when it differs, and None when no such fingerprint exists (including incomplete or manual fits). This comparison ignores row/column order and frequency compression, but does not establish participant identities or independence of samples.

For full-table residuals on complete training or held-out responses, use fit.residuals(data). A held-out residual statistic is descriptive, not a training-sample likelihood-ratio test. For incomplete responses, use the separate incomplete_gof(fit, data) and its saturated-model convergence diagnostics; the report does not invent a missing-data chi-squared table.

AICc convention and limits

from knowledgespaces.estimation import aicc

criterion = aicc(fit.log_likelihood, fit.gof.npar, 1000)
assert criterion == report.AICc
assert report.aicc_sample_convention == "verified_training_respondent_count"

The implemented conventional formula is

[ \mathrm{AICc}=-2\ell+2k+\frac{2k(k+1)}{N-k-1}. ]

The function takes the ordinary log likelihood ℓ. MATLAB’s KST-toolbox stores its negative internally and uses total respondent count for N; after accounting for that sign its displayed formula is the same. This is a source-formula comparison, not a claim of native MATLAB execution. When N ≤ k + 1 the correction is undefined: aicc returns NaN, while the report exports null and an explicit reason, never a negative correction.

The report automatically chooses N only when complete training data are fingerprint-matched and their frequencies are integer-valued. With fractional weights, unknown training provenance, no supplied training data, or an incomplete fit, pass aicc_sample_size explicitly or leave AICc unavailable. An override is a declared training sample-size convention; the size of held-out data is never substituted automatically.

For incomplete data, fit.n_informative excludes wholly missing rows and is the existing BIC convention; fit.n_respondents includes them. Choosing one for AICc must be explicit. Fractional analysis weights are not evidence of an equivalent number of independent observations.

AICc is an algebraic descriptive comparison here, not a generally proved small-sample bias correction for singular or boundary BLIMs. The same qualification applies to nominal-parameter information criteria under nonidentifiability. MD estimates are not maximum-likelihood estimates; their reported full-model likelihood criteria are descriptive rather than an assertion that the usual likelihood-based selection theory applies.

Trade-offs and exports

local = blim_report(fit, include_tradeoffs=True)
assert len(local.tradeoffs.blocks) == 7
assert local.to_dict()["tradeoffs"]["parameters"][0]["family"] == "beta"

Optional trade-offs use the existing blim_tradeoffs calculation. The export records labeled free coordinates, the full-response Jacobian and all seven beta/eta/pi block ranks, tolerances and null-space bases. This is a local derivative calculation. A null direction at one point does not prove global nonidentifiability, and at a boundary it may not be a feasible probability direction. For incomplete fits it concerns the full response model, not the experiment’s observed-mask derivative.

For SLM, include_tradeoffs=True raises because the free state-mass BLIM coordinates are the wrong parameterization. Use slm_fit.jacobian() for that model’s beta/eta/g derivative instead. Expensive complete-pattern enumeration is therefore opt-in and memory-guarded.

References: pks 0.7-0, blim and getMD source definitions for distance, mastery and expected errors; Heller and Wickelmaier (2013), Minimum discrepancy estimation in probabilistic knowledge structures; Stefanutti et al. (2012) for BLIM trade-offs; KST-toolbox blim.m, revision 8b53ecd18f57ec5b8788522d8ebaa12019082330, for the conventional AICc formula.