Reference workflows through explicit compositions¶
The release integrates capabilities rather than duplicating every external
wrapper. cookbook/19_reference_compositions.py is an executable example of
the workflows below. The reference versions are kstMatrix 3.0-0, CDSS 0.3-1
and the pinned KST-toolbox MATLAB revision
8b53ecd18f57ec5b8788522d8ebaa12019082330.
Complete-data ML with parameter-change stopping¶
estimate_blim uses an absolute log-likelihood change for ML convergence.
If the desired control is maximum absolute parameter change, complete data
can use the observed-data estimator without inserting any missing values:
import numpy as np
from knowledgespaces import KnowledgeStructure
from knowledgespaces.estimation import (
BLIMConstraints, IncompleteResponseMatrix, estimate_blim_incomplete,
)
structure = KnowledgeStructure("a", [])
data = IncompleteResponseMatrix(("a",), np.array([[0.0], [1.0]]), np.array([20, 80]))
fixed = BLIMConstraints(beta_fixed={"a": 0}, eta_fixed={"a": 0})
fit = estimate_blim_incomplete(structure, data, constraints=fixed, tol=0.1)
assert fit.converged and fit.n_iterations == 2
np.testing.assert_allclose(fit.pi, [0.2, 0.8])
The stopping quantity is the maximum change across beta, eta and state
probabilities. The example has an analytical ML solution: its first update
changes the prior by 0.3; the second changes it by zero. The same response
likelihood is being maximized when every cell is observed. With ordinary
ResponseMatrix data, construct IncompleteResponseMatrix(tuple(data.items), data.patterns, data.counts) to preserve weights and item identities.
MATLAB blim.m compares this quantity with <= options.xtol; Python uses
< tol. Initialization, clipping and floating-point details also differ,
so this is a stopping-policy correspondence, not an assertion of identical
iteration histories or native MATLAB execution. The MATLAB wrapper’s
options.tol spelling does not redefine its callee’s xtol. Constraints
are explicit: MATLAB’s beta0/eta0 flags indicate fixed zero versus free
errors, not initial error magnitudes. Python uses BLIMConstraints for
fixed values/equality groups and separate initialization arguments.
The result is an IncompleteBLIMEstimate, retaining its observed-likelihood
diagnostics. log_likelihood is the log-likelihood itself (usually a
negative number), whereas the MATLAB loglike field is
the negative log-likelihood. Do not compare them without changing sign.
With missing responses, MATLAB additionally includes an empirical
mask-frequency factor. The Python observed likelihood omits that factor;
under independent masks its contribution is constant in the response
parameters. Neither convention justifies generic missing-data chi-square
degrees of freedom. See bootstrap assumptions.
Smallest quasiordinal completion¶
The kmqspace family capability is a completion, not just a predicate:
from knowledgespaces import SetFamily, SurmiseRelation
family = SetFamily("abc", ["a", "b", "ac", "bc"])
completion = family.union_closure().intersection_closure().to_knowledge_structure()
assert completion.is_quasi_ordinal and completion.n_states == 8
assert family.sets <= completion.states
relation = SurmiseRelation(
family.domain,
[(a, b) for a in family.domain for b in family.domain
if a != b and all(b not in state or a in state for state in family)],
)
assert KnowledgeStructure.from_surmise_relation(relation) == completion
For finite set families, closing under unions and then intersections gives
both closures: distributivity ensures that unions of intersections of the
union-closed family are again intersections of members. Both nullary
operations are included, giving ∅ and Q. The result is the smallest family
with these properties containing the input. Domains must be nonempty to
convert to KnowledgeStructure; SetFamily itself also supports the empty
domain. Guards on the closure methods bound output cardinality.
For a surmise function, start from
function.to_knowledge_base().as_family() and apply the same composition.
For a surmise relation, its downset structure via
KnowledgeStructure.from_surmise_relation(relation) is already
quasiordinal. No maximal powerset is needed as an intermediate object,
although the result itself can be exponentially large.
Raw curriculum relations before closure¶
The raw CDSS skill/object relations apply to a single skill assignment with one teacher per skill. A single-clause attribution is a suitable intermediate representation. Multiple alternative clauses must not be flattened into a conjunctive prerequisite relation:
from knowledgespaces import Attribution, LearningObject, SkillAssignment
def raw_single_clause_relation(attribution: Attribution) -> SurmiseRelation:
if any(len(attribution.clauses_for(q)) != 1 for q in attribution.domain):
raise ValueError("A raw conjunctive relation requires one clause per item.")
return SurmiseRelation(
attribution.domain,
[(a, b) for b in attribution.domain
for a in next(iter(attribution.clauses_for(b))) if a != b],
)
course = SkillAssignment([
LearningObject("intro", frozenset("ab")),
LearningObject("intermediate", frozenset("c"), frozenset("b")),
LearningObject("advanced", frozenset("d"), frozenset("c")),
])
assert course.has_unique_teachers
raw_skills = raw_single_clause_relation(course.skill_attribution())
raw_objects = raw_single_clause_relation(course.object_attribution())
assert raw_skills.relations == {("a", "b"), ("b", "a"), ("b", "c"), ("c", "d")}
assert raw_objects.relations == {("intro", "intermediate"), ("intermediate", "advanced")}
assert ("intro", "advanced") not in raw_objects.relations
assert raw_objects.transitive_closure() == course.derive().object_relation
Pairs point from prerequisite to dependent. Reflexivity is implicit in
SurmiseRelation, while to_matrix() includes its diagonal. The raw graph
can be nontransitive; only transitive_closure() performs that step. A
closed relation was already available in course.derive(); the explicit
composition above additionally exposes the raw graph without changing it.
Other lightweight reference primitives are ordinary Python operations:
A ^ B, len(A ^ B), membership and A <= B for sets. The exact
kmdoubleequal convention is abs(x - y) < tolerance, an absolute strict
comparison. It is not NumPy’s default isclose, which also has a relative
tolerance. The cookbook includes a boundary example.