Adaptive Assessment¶
Deterministic half-split assessment¶
from knowledgespaces import BLIM, BLIMParams, KnowledgeStructure, StatePosterior
from knowledgespaces.assessment import run_assessment
structure = KnowledgeStructure(["a", "b"], [set(), {"a"}, {"a", "b"}])
initial = StatePosterior.uniform(BLIM(structure, BLIMParams(beta=0.0, eta=0.0)))
answers = {"a": True, "b": False}
result = run_assessment(initial, answers.__getitem__, selection="half_split",
threshold=1.0, max_questions=2)
assert result.state == frozenset({"a"}) and result.converged
With zero errors and a uniform prior this is the deterministic candidate
elimination workflow of kst::kassess: each answer removes inconsistent
states. The session stops once one candidate remains, and need not ask every
item. Thus a response pattern outside the structure can still identify a
state compatible with the administered answers, without detecting a conflict
on an unasked item. Use complete-response validation when that is the question.
Native comparison covers every three-item structure and every binary response:
all 320 cases where the response is a model state agree. Off-model outcomes
can differ with question ordering: R compares integer candidate counts;
Python compares floating posterior half-split scores and breaks exact ties
by item order. A mathematical tie can differ by floating roundoff. Neither
workflow turns those off-model outputs into evidence of a noiseless true state.
max_questions and exhausted-item stops remain explicit; incompatible
zero-probability evidence raises instead of silently resetting the posterior.
The BLIM model¶
The Basic Local Independence Model (BLIM) defines the probability of a response given a knowledge state, using two parameters:
\(\beta\) (slip): \(P(\text{incorrect} \mid \text{item mastered})\)
\(\eta\) (guess): \(P(\text{correct} \mid \text{item not mastered})\)
The informative-item condition \(\beta_q + \eta_q < 1\) is checked per
item; supplied parameters that violate it are rejected.
Both parameters can be passed as a scalar (same value for all items) or as
a dict ({item: value}) for per-item values.
Items vs instances¶
In KST/ALEKS, the distinction between items and instances is fundamental:
An item is a problem type / competency (the latent level).
An instance is a concrete question (the observed level).
Each item can have multiple instances of equivalent difficulty. The BLIM update happens at the item level; the engine selects and presents instances to avoid repeating the same question.
One-shot assessment¶
For batch assessment from known responses:
import knowledgespaces as ks
structure = ks.space_from_prerequisites(
["add", "sub", "mul"],
[("add", "sub"), ("sub", "mul")],
)
# Single observation per item
result = ks.assess(structure, {"add": True, "sub": True, "mul": False})
# Multiple observations (different instances of the same item)
result = ks.assess(structure, [
("add", True), ("add", True), # two instances of addition
("sub", True),
("mul", False),
])
print(result["state"]) # most likely knowledge state
print(result["probability"]) # posterior probability
print(result["mastery"]) # per-item mastery probabilities
print(result["outer_fringe"]) # what to learn next
print(result["inner_fringe"]) # most recently consolidated
Both assess and adaptive_assess accept an optional prior over states.
To reuse parameters and prior estimated via EM, see assess_from_fit and
adaptive_assess_from_fit in estimation.
Adaptive assessment¶
The engine selects the most informative question using Expected Information Gain (EIG):
where \(H\) is Shannon entropy and the expectation is over both possible responses.
With instances (recommended)¶
result = ks.adaptive_assess(
structure,
ask_fn=lambda instance_id: ask_student(instance_id),
instances={
"add": ["3+2", "7+5", "12+9"],
"sub": ["8-3", "15-7"],
"mul": ["4*3", "6*7"],
},
beta=0.1,
eta=0.2,
threshold=0.85,
max_questions=15,
)
Simple mode (one question per item)¶
result = ks.adaptive_assess(
structure, lambda item: item in {"add", "sub"}
)
Low-level control¶
For full control over the assessment loop:
from knowledgespaces import BLIM, BLIMParams, StatePosterior
from knowledgespaces.assessment import select_item_eig, is_converged
blim = BLIM(structure, BLIMParams(beta=0.1, eta=0.2))
posterior = StatePosterior.uniform(blim)
posterior = posterior.update("add", True)
posterior = posterior.update("sub", True)
print(posterior.entropy)
print(posterior.most_likely_state)
print(posterior.marginal_mastery())
The half-split rule¶
Alongside EIG, the classic half-split questioning rule of knowledge space theory is available (Falmagne & Doignon, 1988; Falmagne & Doignon, 2011, ch. 13): pick the item whose current mastery probability \(p_q = P(q \in K)\) is closest to \(1/2\) — the item that best bisects the probability mass over plausible states.
from knowledgespaces.assessment import select_item_half_split
best = select_item_half_split(posterior, exclude={"add"})
print(best.item) # the best bisecting item
print(best.score) # min(p, 1 - p) in [0, 0.5]; 0.5 = perfect bisection
Unlike EIG, half-split ignores the error parameters \(\beta, \eta\) at selection time — it depends only on the current state distribution. Under noise-free responding the two rules coincide; with noise they can diverge, which makes half-split the natural baseline when evaluating EIG-based selection.
The rule is also available end to end, without writing the loop yourself:
adaptive_assess(structure, ask_fn, selection="half_split") in the
high-level API, and ks assess --selection half-split on the command
line (item-level mode).
Instance-level selection¶
from knowledgespaces.assessment import InstancePool, select_instance_eig
pool = InstancePool.from_dict({
"add": ["add_q1", "add_q2"],
"sub": ["sub_q1", "sub_q2"],
"mul": ["mul_q1"],
})
best = select_instance_eig(posterior, pool, asked={"add_q1"})
print(best.instance_id) # e.g. "sub_q1"
print(best.item) # "sub"
print(best.score) # EIG value
The canonical informative rule¶
select_item_informative implements Definition 13.4.8 and Equation (13.14)
of Falmagne & Doignon’s Learning Spaces (2011):
Choose an item minimizing this quantity. The returned score is
\(H(\pi)-\widetilde H(q,\pi)\), in bits; it is not the Bayesian EIG.
The latter uses the BLIM predictive probability
\(P(R_q=1)=(1-\beta_q)p_q+\eta_q(1-p_q)\) as the weight. With noisy responses,
the rules can select different items. kstMatrix::kmassessinformative
uses the canonical mastery weights. This is a theoretical convention,
not an implementation error.
Both the half-split and informative definitions draw uniformly among
maximizers. Supply rng=random.Random(seed) to the item selectors or
run_assessment for uniform exact computed score ties. Without rng,
they retain deterministic candidate-order ties for compatibility. Duplicate
candidate labels do not change the sampling weights. Floating-point scores
that differ are not treated as exact ties. The existing instance EIG helper
retains its numpy.isclose tie convention.
Complete sessions and multiplicative updating¶
import random
from knowledgespaces.assessment import MultiplicativePosterior, run_assessment
initial = MultiplicativePosterior.uniform(structure, zeta0=4, zeta1=5)
session = run_assessment(
initial,
lambda item: item in {"add", "sub"},
selection="informative", # or "half_split"
threshold=0.85,
max_questions=15,
rng=random.Random(42),
track_probabilities=True,
)
print(session.state, session.probability, session.stop_reason)
print(session.questions_asked, session.steps)
The multiplicative update rewards agreement by \(\zeta_{q,r}>1\) and
normalizes, as in Equations (13.9)–(13.10). A StatePosterior instead uses
its BLIM; it also supports selection="eig". The supplied prior is not
changed. Every recorded probability vector follows session.posterior.states.
When requested, the history includes the initial vector and every update;
otherwise it is empty to save memory. Selection and update times exclude
waiting for the response callback.
Stopping means one of:
"threshold": the largest state mass is at least the threshold;"exhausted": all available items or instances have been used;"max_questions": the question limit has been reached.
The threshold has priority, then exhaustion, then the question limit.
A prior already meeting the threshold produces a zero-question session.
session.converged reports threshold attainment only; it does not prove
correct identification or calibrated coverage. Nonconverged sessions keep
all their results. The high-level adaptive_assess dictionary also exposes
converged, stop_reason and posterior.
All three policies work with InstancePool in run_assessment, and with
instances=... in adaptive_assess. Each instance is administered at most
once. Without a pool, each item is used at most once unless
repeat_items=True. That option invokes the callback again, and its use
with Bayesian updating assumes fresh, conditionally independent responses
to a fixed latent state. The CLI currently offers EIG and half-split at item
level and EIG with instances.
Responses must be booleans or numeric 0/1. Strings, None, missing values
and other numbers are rejected. Callback and impossible-evidence errors
propagate; they are not turned into a successful or silently discarded run.
Evaluate observed responses versus simulate latent states¶
evaluate_assessment(initial, response_matrix, ...) replays one session
per stored complete binary row, aligns item labels, and retains row weights.
Each row starts from the supplied initial distribution. Its distances are:
Output |
Meaning |
|---|---|
|
Hamming distance between diagnosed state and full observed response pattern |
|
Minimum Hamming distance from that pattern to any state in the structure |
|
Difference of the two preceding distances |
These outputs parallel the response-based quantities called assessment error,
distance and net assessment error in kstMatrix. They do not measure
accuracy against a known latent state. The three weighted means,
mean_questions and convergence_rate, include all sessions. An aggregated
row is evaluated once: its weight does not create independent randomized
tie decisions for every person it represents. Use individual rows to study
those decisions independently.
For compatibility experiments, repeat_items=True reuses the same stored
answer if the policy selects an item again. This matches fixed-pattern
replay in kstMatrix; it is not new independent Bayesian evidence. The
Python default avoids reusing stored answers. R’s kmassess stops at a
strict > threshold and returns NULL after its fixed limit of twice the
item count; Python uses >=, an explicit question limit and retains results.
R’s informative selector can return several indices on a tie; Python returns
one item according to the documented tie policy. Neither R’s plotting side
effects nor its clock timings are numerical compatibility targets.
To measure state recovery under a response-generating model, use
simulate_assessment instead:
from knowledgespaces import BLIM, BLIMParams, StatePosterior
from knowledgespaces.assessment import simulate_assessment
model = BLIM(structure, BLIMParams(beta=0.1, eta=0.2))
initial = StatePosterior.uniform(model)
evaluation = simulate_assessment(
initial,
[frozenset({"add", "sub"})] * 100,
response_model=model,
selection="eig",
repeat_items=True,
max_questions=20,
seed=42,
)
print(evaluation.mean_state_distance, evaluation.exact_state_rate)
print(evaluation.mean_questions, evaluation.convergence_rate)
Each supplied true state identifies one simulated person. Every administered
response is a new Bernoulli draw, including repeated questions. The generating
and assessing models may differ but must share item labels; true states must
belong to the generating structure. Simulation reports state_distances
and exact_state_rate; observed-response distances are None, since no
complete response vector is generated. A seed fixes local response and
question-selection generators. Run cookbook/14_assessment_workflows.py
for a complete example using observed probability data and simulated states.
The workflows do not establish asymptotic identification for every combination. Chapter 13’s results have additional hypotheses; §13.6 explicitly distinguishes the proved half-split results from the unresolved multiplicative/informative combination in that source. A finite question budget or reused fixed answers also differs from an indefinitely continuing stochastic response process.