Bundled Datasets

The knowledgespaces.datasets module ships classical example data with documented provenance, ready for estimation, simulation, and assessment examples.

The 0.2.0 selection has three sources and four loaders: the fictitious Doignon–Falmagne example for reproducible numerical checks, probability for the main empirical application, and PISA for IITA. Functional coverage of a reference package does not mean that every dataset it ships is redistributed here. All 21 dataset objects and 14 auxiliary example files in the pinned comparison inventory have an explicit release decision.

Doignon & Falmagne (1999, chapter 7)

The five-item example paired with the response patterns of 1000 fictitious respondents, distributed as DoignonFalmagne7 in the R package pks:

from knowledgespaces.datasets import load_doignon_falmagne_7

ds = load_doignon_falmagne_7()
ds.structure.n_states            # 9 states on items a..e
ds.structure.is_learning_space   # True
ds.data                          # ResponseMatrix (pattern frequencies, N = 1000)
ds.source                        # bibliographic provenance

The structure/frequency loaders return a KSTDataset with the fields structure (a KnowledgeStructure), data (a ResponseMatrix), name, description, and source.

Using a dataset

The structure and data plug directly into the estimation layer:

from knowledgespaces.estimation import estimate_blim

fit = estimate_blim(ds.structure, ds.data)
fit.log_likelihood

The same dataset anchors the cross-validation against pks in tests/test_cross_validation.py, so the numbers it produces are pinned in continuous integration.

The numbers originate in the cited book; pks 0.7-0 is the distribution used for numerical cross-checking and declares GPL (>= 2) for its package. This is a fictitious textbook example, not a sample of observed participants.

Probability: Anselmi & Wickelmaier

The probability study is available offline through the installed package:

from knowledgespaces.datasets import load_probability
from knowledgespaces.estimation import estimate_blim

ds = load_probability(wave="pre", model="K2")
assert ds.data.n_respondents == 345
assert ds.structure.n_states == 13
fit = estimate_blim(ds.structure, ds.data, method="MD")
print(fit.log_likelihood)  # approximately -1545.6976905728

Choose wave="pre" or "post", and model="K1" (16 states, conjunctive sf1) or "K2" (13 states, alternative competencies sf2). ds.skill_map exposes the delineating skill function; it uses four skills (cp, id, pb, un) and item labels i01 through i12. The structures are candidate models: neither is union-closed and neither is an observed true knowledge structure.

The 345 completers are selected from the original 504 cases using the source example’s !is.na(probability$b201) rule. There are 110 pre and 84 post response patterns. Both waves concern the same people, but aggregate frequencies do not preserve their pairing. Do not pair rows or treat expanded frequencies as a longitudinal individual dataset. The loader does not provide group assignments for treatment comparisons.

Source: Anselmi and Wickelmaier, data collected in Tuebingen in 2010, probability in pks 0.7-0. The data retain GPL-2.0-or-later terms, separately from the MIT Python code. Source metadata, the full license, CSV checksums and an R export script are packaged with the data. The source export reproduces the two frequency tables and K1/K2 reference structures byte for byte.

See cookbook/03_probability.py for the four dataset selections. The paper replication archive contains fuller estimation and IITA comparisons.

Individual probability records and pairing

from knowledgespaces.datasets import load_probability_individual

individual = load_probability_individual()  # all 504 retained source cases
paired = load_probability_individual(sample="completers")  # same 345 IDs in both waves
assert individual.post.n_informative == 345
assert len(paired.case_ids) == 345
pre = paired.pre.to_complete()
post = paired.post.to_complete()

pre and post in the returned dataset are immutable IncompleteResponseMatrix objects, row-aligned with case_ids and records. Every record retains all 68 source columns: original factor labels, numeric values, UTC timestamp strings, and None for missing fields. schema() returns source types and factor levels; source_documentation() returns the original pks Rd codebook and item wordings. raw_answers("pre"|"post") returns the numeric p responses, not their dichotomous scores.

The source binary pre scores are complete for all 504 cases. The 159 non-completers have all twelve post binary scores missing. The source already codes 417 omitted numeric pre responses as binary zero; the loader preserves that distinction and does not silently rescore them as NaN. Raw numeric post omissions and source-excluded post score rows are likewise different kinds of information. The p110 wording/correct answer differs between lab and online; the original codebook records this.

Selections mode="lab"|"online" and condition="basic"|"enhan" can be combined with sample. They preserve original IDs/order and filter both waves together. There are 26 lab cases, all enhanced, and 478 online cases (221 enhanced, 257 basic); empty selections raise. Among completers there are 26 lab-enhanced, 155 online-enhanced and 164 online-basic cases. These counts do not by themselves justify a causal treatment analysis or ignore selection/dropout. All 504 rows are the source package’s retained sample, not the entire originally recruited population.

The individual CSV, schema, codebook and export script retain GPL-2.0-or-later terms. Source numeric values are checked against R doubles, timestamps against their original instants, and completer frequencies against the existing aggregate loader. The incomplete-data guide explains fitting assumptions; cookbook/07_incomplete_probability.py runs the paired-data and post-only examples offline.

PISA 2003: DAKS response data

from knowledgespaces.datasets import load_pisa
from knowledgespaces.derivation import iita

pisa = load_pisa()
assert pisa.data.patterns.shape == (340, 5)
fit = iita(pisa.data, version="minimized")
frequencies = load_pisa(aggregate=True).data

This is the complete five-item response data frame from DAKS 2.1-3: 340 German students in PISA 2003, with the source’s dichotomized scores. The default preserves every original row and labels a..e. The aggregate view returns identical-pattern counts, preserving analysis totals. ResponseDataset exposes data, name, description, source and source_license. It supplies no ground-truth structure: the example’s IITA relation is estimated from these responses.

The source does not supply the item wording, original PISA item identifiers, original polytomous scores or dichotomization rules. The CSV’s source_row values preserve R row positions, not recovered participant identifiers. There is no filtering, imputation or new dichotomization in this export. Do not infer item content or reconstruct scores that are absent from the source. Loading requires neither R nor internet access.

GPL-2.0-or-later data terms, the source codebook and package metadata are packaged under knowledgespaces/datasets/data/pisa/, separately from MIT Python code. All CSV values have been compared with the DAKS source archive; loader/IITA tests also compare fresh native R counts and all three criteria. See the frequency and quotient guide for descriptive reports.

Selection from the comparison packages

The remaining resources are deliberately outside the bundled data selection. They can motivate later applications without being prerequisites for the algorithms or for the probability case study.

Source and objects

Decision for 0.2.0

Reason

pks: circles, angles

Not bundled

Additional geometry pre/post examples; the chosen probability study already supplies paired records. The geometry export omits its held-out sample.

pks: chess

Not bundled

Additional 16-item application with corrected source relations.

pks: density97, matter97

Not bundled

Additional science examples whose frequencies were reconstructed from published histograms.

pks: endm

Not bundled

Additional artificial estimation example; numerical checks already use the textbook and independent synthetic inputs.

pks: subtraction13, fraction17

Not bundled

Additional school arithmetic studies. Any later fraction loader must resolve the codebook’s scoring convention explicitly.

kstMatrix: cad, fractions, phsg, readwrite

Not bundled

Additional expert-derived structures and application content, not missing structural APIs.

kstMatrix: xpl; CbKST: exampledata

Upstream examples acknowledged

Independently authored examples demonstrate these workflows; their named data objects are not copied.

kstpy: data.xpl_basis

Duplicate component, not bundled

Its set family equals kstMatrix::xpl$basis; the other xpl components are separate resources.

MATLAB: dataSD.mat, data_fraction.mat, skill_functionSD.mat

Not bundled

Additional application data and skill maps are unnecessary for the chosen case. No numerical identity with the pks arithmetic data is asserted.

CbKST: four ODS examples; CDSS: ten CSV/XLSX/ODS examples

Upstream examples acknowledged

These teach import, competence mapping and course-skill workflows. Original synthetic fixtures exercise the supported formats and methods without redistributing the course content.

These are scope decisions, not assertions that omitted resources cannot be redistributed. The pinned pks and DAKS packages declare GPL version 2 or later; kstMatrix, CbKST and CDSS declare GPL version 3. The MATLAB repository declares CC BY-NC-SA 4.0, and no separate dataset-specific grant was established in the inspected metadata. kstpy’s package metadata says LGPL-3.0-or-later, while its README and LICENSE say GPL version 3; that inconsistency remains explicit rather than being resolved by assumption. None of those omitted resources is included merely because a compatible Python method exists.

The release checks compare all included probability and PISA exports against their pinned native sources, validate the packaged checksum manifests, and check the textbook frequencies against pks. Exporting the same source again validates transcription and transformations; it does not establish that the study sample is representative or that a fitted model is scientifically true.