Reference workflows through explicit compositions

The release integrates capabilities rather than duplicating every external wrapper. cookbook/19_reference_compositions.py is an executable example of the workflows below. The reference versions are kstMatrix 3.0-0, CDSS 0.3-1 and the pinned KST-toolbox MATLAB revision 8b53ecd18f57ec5b8788522d8ebaa12019082330.

Complete-data ML with parameter-change stopping

estimate_blim uses an absolute log-likelihood change for ML convergence. If the desired control is maximum absolute parameter change, complete data can use the observed-data estimator without inserting any missing values:

import numpy as np
from knowledgespaces import KnowledgeStructure
from knowledgespaces.estimation import (
    BLIMConstraints, IncompleteResponseMatrix, estimate_blim_incomplete,
)

structure = KnowledgeStructure("a", [])
data = IncompleteResponseMatrix(("a",), np.array([[0.0], [1.0]]), np.array([20, 80]))
fixed = BLIMConstraints(beta_fixed={"a": 0}, eta_fixed={"a": 0})
fit = estimate_blim_incomplete(structure, data, constraints=fixed, tol=0.1)
assert fit.converged and fit.n_iterations == 2
np.testing.assert_allclose(fit.pi, [0.2, 0.8])

The stopping quantity is the maximum change across beta, eta and state probabilities. The example has an analytical ML solution: its first update changes the prior by 0.3; the second changes it by zero. The same response likelihood is being maximized when every cell is observed. With ordinary ResponseMatrix data, construct IncompleteResponseMatrix(tuple(data.items), data.patterns, data.counts) to preserve weights and item identities.

MATLAB blim.m compares this quantity with <= options.xtol; Python uses < tol. Initialization, clipping and floating-point details also differ, so this is a stopping-policy correspondence, not an assertion of identical iteration histories or native MATLAB execution. The MATLAB wrapper’s options.tol spelling does not redefine its callee’s xtol. Constraints are explicit: MATLAB’s beta0/eta0 flags indicate fixed zero versus free errors, not initial error magnitudes. Python uses BLIMConstraints for fixed values/equality groups and separate initialization arguments.

The result is an IncompleteBLIMEstimate, retaining its observed-likelihood diagnostics. log_likelihood is the log-likelihood itself (usually a negative number), whereas the MATLAB loglike field is the negative log-likelihood. Do not compare them without changing sign. With missing responses, MATLAB additionally includes an empirical mask-frequency factor. The Python observed likelihood omits that factor; under independent masks its contribution is constant in the response parameters. Neither convention justifies generic missing-data chi-square degrees of freedom. See bootstrap assumptions.

Smallest quasiordinal completion

The kmqspace family capability is a completion, not just a predicate:

from knowledgespaces import SetFamily, SurmiseRelation

family = SetFamily("abc", ["a", "b", "ac", "bc"])
completion = family.union_closure().intersection_closure().to_knowledge_structure()
assert completion.is_quasi_ordinal and completion.n_states == 8
assert family.sets <= completion.states

relation = SurmiseRelation(
    family.domain,
    [(a, b) for a in family.domain for b in family.domain
     if a != b and all(b not in state or a in state for state in family)],
)
assert KnowledgeStructure.from_surmise_relation(relation) == completion

For finite set families, closing under unions and then intersections gives both closures: distributivity ensures that unions of intersections of the union-closed family are again intersections of members. Both nullary operations are included, giving ∅ and Q. The result is the smallest family with these properties containing the input. Domains must be nonempty to convert to KnowledgeStructure; SetFamily itself also supports the empty domain. Guards on the closure methods bound output cardinality.

For a surmise function, start from function.to_knowledge_base().as_family() and apply the same composition. For a surmise relation, its downset structure via KnowledgeStructure.from_surmise_relation(relation) is already quasiordinal. No maximal powerset is needed as an intermediate object, although the result itself can be exponentially large.

Raw curriculum relations before closure

The raw CDSS skill/object relations apply to a single skill assignment with one teacher per skill. A single-clause attribution is a suitable intermediate representation. Multiple alternative clauses must not be flattened into a conjunctive prerequisite relation:

from knowledgespaces import Attribution, LearningObject, SkillAssignment

def raw_single_clause_relation(attribution: Attribution) -> SurmiseRelation:
    if any(len(attribution.clauses_for(q)) != 1 for q in attribution.domain):
        raise ValueError("A raw conjunctive relation requires one clause per item.")
    return SurmiseRelation(
        attribution.domain,
        [(a, b) for b in attribution.domain
         for a in next(iter(attribution.clauses_for(b))) if a != b],
    )

course = SkillAssignment([
    LearningObject("intro", frozenset("ab")),
    LearningObject("intermediate", frozenset("c"), frozenset("b")),
    LearningObject("advanced", frozenset("d"), frozenset("c")),
])
assert course.has_unique_teachers
raw_skills = raw_single_clause_relation(course.skill_attribution())
raw_objects = raw_single_clause_relation(course.object_attribution())
assert raw_skills.relations == {("a", "b"), ("b", "a"), ("b", "c"), ("c", "d")}
assert raw_objects.relations == {("intro", "intermediate"), ("intermediate", "advanced")}
assert ("intro", "advanced") not in raw_objects.relations
assert raw_objects.transitive_closure() == course.derive().object_relation

Pairs point from prerequisite to dependent. Reflexivity is implicit in SurmiseRelation, while to_matrix() includes its diagonal. The raw graph can be nontransitive; only transitive_closure() performs that step. A closed relation was already available in course.derive(); the explicit composition above additionally exposes the raw graph without changing it.

Other lightweight reference primitives are ordinary Python operations: A ^ B, len(A ^ B), membership and A <= B for sets. The exact kmdoubleequal convention is abs(x - y) < tolerance, an absolute strict comparison. It is not NumPy’s default isclose, which also has a relative tolerance. The cookbook includes a boundary example.