Adaptive Assessment

Deterministic half-split assessment

from knowledgespaces import BLIM, BLIMParams, KnowledgeStructure, StatePosterior
from knowledgespaces.assessment import run_assessment

structure = KnowledgeStructure(["a", "b"], [set(), {"a"}, {"a", "b"}])
initial = StatePosterior.uniform(BLIM(structure, BLIMParams(beta=0.0, eta=0.0)))
answers = {"a": True, "b": False}
result = run_assessment(initial, answers.__getitem__, selection="half_split",
                        threshold=1.0, max_questions=2)
assert result.state == frozenset({"a"}) and result.converged

With zero errors and a uniform prior this is the deterministic candidate elimination workflow of kst::kassess: each answer removes inconsistent states. The session stops once one candidate remains, and need not ask every item. Thus a response pattern outside the structure can still identify a state compatible with the administered answers, without detecting a conflict on an unasked item. Use complete-response validation when that is the question.

Native comparison covers every three-item structure and every binary response: all 320 cases where the response is a model state agree. Off-model outcomes can differ with question ordering: R compares integer candidate counts; Python compares floating posterior half-split scores and breaks exact ties by item order. A mathematical tie can differ by floating roundoff. Neither workflow turns those off-model outputs into evidence of a noiseless true state. max_questions and exhausted-item stops remain explicit; incompatible zero-probability evidence raises instead of silently resetting the posterior.

The BLIM model

The Basic Local Independence Model (BLIM) defines the probability of a response given a knowledge state, using two parameters:

  • \(\beta\) (slip): \(P(\text{incorrect} \mid \text{item mastered})\)

  • \(\eta\) (guess): \(P(\text{correct} \mid \text{item not mastered})\)

The informative-item condition \(\beta_q + \eta_q < 1\) is checked per item; supplied parameters that violate it are rejected. Both parameters can be passed as a scalar (same value for all items) or as a dict ({item: value}) for per-item values.

Items vs instances

In KST/ALEKS, the distinction between items and instances is fundamental:

  • An item is a problem type / competency (the latent level).

  • An instance is a concrete question (the observed level).

Each item can have multiple instances of equivalent difficulty. The BLIM update happens at the item level; the engine selects and presents instances to avoid repeating the same question.

One-shot assessment

For batch assessment from known responses:

import knowledgespaces as ks

structure = ks.space_from_prerequisites(
    ["add", "sub", "mul"],
    [("add", "sub"), ("sub", "mul")],
)

# Single observation per item
result = ks.assess(structure, {"add": True, "sub": True, "mul": False})

# Multiple observations (different instances of the same item)
result = ks.assess(structure, [
    ("add", True), ("add", True),   # two instances of addition
    ("sub", True),
    ("mul", False),
])

print(result["state"])        # most likely knowledge state
print(result["probability"])  # posterior probability
print(result["mastery"])      # per-item mastery probabilities
print(result["outer_fringe"]) # what to learn next
print(result["inner_fringe"]) # most recently consolidated

Both assess and adaptive_assess accept an optional prior over states. To reuse parameters and prior estimated via EM, see assess_from_fit and adaptive_assess_from_fit in estimation.

Adaptive assessment

The engine selects the most informative question using Expected Information Gain (EIG):

\[\text{EIG}(q) = H(\pi) - \mathbb{E}[H(\pi \mid R_q)]\]

where \(H\) is Shannon entropy and the expectation is over both possible responses.

Simple mode (one question per item)

result = ks.adaptive_assess(
    structure, lambda item: item in {"add", "sub"}
)

Low-level control

For full control over the assessment loop:

from knowledgespaces import BLIM, BLIMParams, StatePosterior
from knowledgespaces.assessment import select_item_eig, is_converged

blim = BLIM(structure, BLIMParams(beta=0.1, eta=0.2))
posterior = StatePosterior.uniform(blim)

posterior = posterior.update("add", True)
posterior = posterior.update("sub", True)

print(posterior.entropy)
print(posterior.most_likely_state)
print(posterior.marginal_mastery())

The half-split rule

Alongside EIG, the classic half-split questioning rule of knowledge space theory is available (Falmagne & Doignon, 1988; Falmagne & Doignon, 2011, ch. 13): pick the item whose current mastery probability \(p_q = P(q \in K)\) is closest to \(1/2\) — the item that best bisects the probability mass over plausible states.

from knowledgespaces.assessment import select_item_half_split

best = select_item_half_split(posterior, exclude={"add"})
print(best.item)   # the best bisecting item
print(best.score)  # min(p, 1 - p) in [0, 0.5]; 0.5 = perfect bisection

Unlike EIG, half-split ignores the error parameters \(\beta, \eta\) at selection time — it depends only on the current state distribution. Under noise-free responding the two rules coincide; with noise they can diverge, which makes half-split the natural baseline when evaluating EIG-based selection.

The rule is also available end to end, without writing the loop yourself: adaptive_assess(structure, ask_fn, selection="half_split") in the high-level API, and ks assess --selection half-split on the command line (item-level mode).

Instance-level selection

from knowledgespaces.assessment import InstancePool, select_instance_eig

pool = InstancePool.from_dict({
    "add": ["add_q1", "add_q2"],
    "sub": ["sub_q1", "sub_q2"],
    "mul": ["mul_q1"],
})

best = select_instance_eig(posterior, pool, asked={"add_q1"})
print(best.instance_id)  # e.g. "sub_q1"
print(best.item)         # "sub"
print(best.score)        # EIG value

The canonical informative rule

select_item_informative implements Definition 13.4.8 and Equation (13.14) of Falmagne & Doignon’s Learning Spaces (2011):

\[ \widetilde H(q,\pi)=p_q H(u(1,q,\pi))+(1-p_q)H(u(0,q,\pi)), \qquad p_q=\sum_{K\ni q}\pi(K). \]

Choose an item minimizing this quantity. The returned score is \(H(\pi)-\widetilde H(q,\pi)\), in bits; it is not the Bayesian EIG. The latter uses the BLIM predictive probability \(P(R_q=1)=(1-\beta_q)p_q+\eta_q(1-p_q)\) as the weight. With noisy responses, the rules can select different items. kstMatrix::kmassessinformative uses the canonical mastery weights. This is a theoretical convention, not an implementation error.

Both the half-split and informative definitions draw uniformly among maximizers. Supply rng=random.Random(seed) to the item selectors or run_assessment for uniform exact computed score ties. Without rng, they retain deterministic candidate-order ties for compatibility. Duplicate candidate labels do not change the sampling weights. Floating-point scores that differ are not treated as exact ties. The existing instance EIG helper retains its numpy.isclose tie convention.

Complete sessions and multiplicative updating

import random
from knowledgespaces.assessment import MultiplicativePosterior, run_assessment

initial = MultiplicativePosterior.uniform(structure, zeta0=4, zeta1=5)
session = run_assessment(
    initial,
    lambda item: item in {"add", "sub"},
    selection="informative",     # or "half_split"
    threshold=0.85,
    max_questions=15,
    rng=random.Random(42),
    track_probabilities=True,
)
print(session.state, session.probability, session.stop_reason)
print(session.questions_asked, session.steps)

The multiplicative update rewards agreement by \(\zeta_{q,r}>1\) and normalizes, as in Equations (13.9)–(13.10). A StatePosterior instead uses its BLIM; it also supports selection="eig". The supplied prior is not changed. Every recorded probability vector follows session.posterior.states. When requested, the history includes the initial vector and every update; otherwise it is empty to save memory. Selection and update times exclude waiting for the response callback.

Stopping means one of:

  • "threshold": the largest state mass is at least the threshold;

  • "exhausted": all available items or instances have been used;

  • "max_questions": the question limit has been reached.

The threshold has priority, then exhaustion, then the question limit. A prior already meeting the threshold produces a zero-question session. session.converged reports threshold attainment only; it does not prove correct identification or calibrated coverage. Nonconverged sessions keep all their results. The high-level adaptive_assess dictionary also exposes converged, stop_reason and posterior.

All three policies work with InstancePool in run_assessment, and with instances=... in adaptive_assess. Each instance is administered at most once. Without a pool, each item is used at most once unless repeat_items=True. That option invokes the callback again, and its use with Bayesian updating assumes fresh, conditionally independent responses to a fixed latent state. The CLI currently offers EIG and half-split at item level and EIG with instances.

Responses must be booleans or numeric 0/1. Strings, None, missing values and other numbers are rejected. Callback and impossible-evidence errors propagate; they are not turned into a successful or silently discarded run.

Evaluate observed responses versus simulate latent states

evaluate_assessment(initial, response_matrix, ...) replays one session per stored complete binary row, aligns item labels, and retains row weights. Each row starts from the supplied initial distribution. Its distances are:

Output

Meaning

response_distances

Hamming distance between diagnosed state and full observed response pattern

structure_distances

Minimum Hamming distance from that pattern to any state in the structure

excess_response_distances

Difference of the two preceding distances

These outputs parallel the response-based quantities called assessment error, distance and net assessment error in kstMatrix. They do not measure accuracy against a known latent state. The three weighted means, mean_questions and convergence_rate, include all sessions. An aggregated row is evaluated once: its weight does not create independent randomized tie decisions for every person it represents. Use individual rows to study those decisions independently.

For compatibility experiments, repeat_items=True reuses the same stored answer if the policy selects an item again. This matches fixed-pattern replay in kstMatrix; it is not new independent Bayesian evidence. The Python default avoids reusing stored answers. R’s kmassess stops at a strict > threshold and returns NULL after its fixed limit of twice the item count; Python uses >=, an explicit question limit and retains results. R’s informative selector can return several indices on a tie; Python returns one item according to the documented tie policy. Neither R’s plotting side effects nor its clock timings are numerical compatibility targets.

To measure state recovery under a response-generating model, use simulate_assessment instead:

from knowledgespaces import BLIM, BLIMParams, StatePosterior
from knowledgespaces.assessment import simulate_assessment

model = BLIM(structure, BLIMParams(beta=0.1, eta=0.2))
initial = StatePosterior.uniform(model)
evaluation = simulate_assessment(
    initial,
    [frozenset({"add", "sub"})] * 100,
    response_model=model,
    selection="eig",
    repeat_items=True,
    max_questions=20,
    seed=42,
)
print(evaluation.mean_state_distance, evaluation.exact_state_rate)
print(evaluation.mean_questions, evaluation.convergence_rate)

Each supplied true state identifies one simulated person. Every administered response is a new Bernoulli draw, including repeated questions. The generating and assessing models may differ but must share item labels; true states must belong to the generating structure. Simulation reports state_distances and exact_state_rate; observed-response distances are None, since no complete response vector is generated. A seed fixes local response and question-selection generators. Run cookbook/14_assessment_workflows.py for a complete example using observed probability data and simulated states.

The workflows do not establish asymptotic identification for every combination. Chapter 13’s results have additional hypotheses; §13.6 explicitly distinguishes the proved half-split results from the unresolved multiplicative/informative combination in that source. A finite question budget or reused fixed answers also differs from an indefinitely continuing stochastic response process.