Bundled Datasets¶
The knowledgespaces.datasets module ships classical example data with
documented provenance, ready for estimation, simulation, and assessment
examples.
The 0.2.0 selection has three sources and four loaders: the fictitious Doignon–Falmagne example for reproducible numerical checks, probability for the main empirical application, and PISA for IITA. Functional coverage of a reference package does not mean that every dataset it ships is redistributed here. All 21 dataset objects and 14 auxiliary example files in the pinned comparison inventory have an explicit release decision.
Doignon & Falmagne (1999, chapter 7)¶
The five-item example paired with the response patterns of 1000
fictitious respondents, distributed as DoignonFalmagne7 in the R
package pks:
from knowledgespaces.datasets import load_doignon_falmagne_7
ds = load_doignon_falmagne_7()
ds.structure.n_states # 9 states on items a..e
ds.structure.is_learning_space # True
ds.data # ResponseMatrix (pattern frequencies, N = 1000)
ds.source # bibliographic provenance
The structure/frequency loaders return a KSTDataset with the fields structure
(a KnowledgeStructure), data (a ResponseMatrix), name,
description, and source.
Using a dataset¶
The structure and data plug directly into the estimation layer:
from knowledgespaces.estimation import estimate_blim
fit = estimate_blim(ds.structure, ds.data)
fit.log_likelihood
The same dataset anchors the cross-validation against pks in
tests/test_cross_validation.py, so the numbers it produces are pinned
in continuous integration.
The numbers originate in the cited book; pks 0.7-0 is the distribution used
for numerical cross-checking and declares GPL (>= 2) for its package.
This is a fictitious textbook example, not a sample of observed participants.
Probability: Anselmi & Wickelmaier¶
The probability study is available offline through the installed package:
from knowledgespaces.datasets import load_probability
from knowledgespaces.estimation import estimate_blim
ds = load_probability(wave="pre", model="K2")
assert ds.data.n_respondents == 345
assert ds.structure.n_states == 13
fit = estimate_blim(ds.structure, ds.data, method="MD")
print(fit.log_likelihood) # approximately -1545.6976905728
Choose wave="pre" or "post", and model="K1" (16 states,
conjunctive sf1) or "K2" (13 states, alternative competencies sf2).
ds.skill_map exposes the delineating skill function; it uses four
skills (cp, id, pb, un) and item labels i01 through i12.
The structures are candidate models: neither is union-closed and
neither is an observed true knowledge structure.
The 345 completers are selected from the original 504 cases using the
source example’s !is.na(probability$b201) rule. There are 110 pre and
84 post response patterns. Both waves concern the same people, but
aggregate frequencies do not preserve their pairing. Do not pair
rows or treat expanded frequencies as a longitudinal individual dataset.
The loader does not provide group assignments for treatment comparisons.
Source: Anselmi and Wickelmaier, data collected in Tuebingen in 2010,
probability in pks 0.7-0.
The data retain GPL-2.0-or-later terms, separately from the MIT
Python code. Source metadata, the full license, CSV checksums and an R
export script are packaged with the data. The source export reproduces
the two frequency tables and K1/K2 reference structures byte for byte.
See cookbook/03_probability.py for the four dataset selections. The
paper replication archive contains fuller estimation and IITA comparisons.
Individual probability records and pairing¶
from knowledgespaces.datasets import load_probability_individual
individual = load_probability_individual() # all 504 retained source cases
paired = load_probability_individual(sample="completers") # same 345 IDs in both waves
assert individual.post.n_informative == 345
assert len(paired.case_ids) == 345
pre = paired.pre.to_complete()
post = paired.post.to_complete()
pre and post in the returned dataset are immutable
IncompleteResponseMatrix objects, row-aligned with case_ids and records.
Every record retains all 68 source columns: original factor labels, numeric
values, UTC timestamp strings, and None for missing fields. schema()
returns source types and factor levels; source_documentation() returns
the original pks Rd codebook and item wordings. raw_answers("pre"|"post")
returns the numeric p responses, not their dichotomous scores.
The source binary pre scores are complete for all 504 cases. The 159 non-completers have all twelve post binary scores missing. The source already codes 417 omitted numeric pre responses as binary zero; the loader preserves that distinction and does not silently rescore them as NaN. Raw numeric post omissions and source-excluded post score rows are likewise different kinds of information. The p110 wording/correct answer differs between lab and online; the original codebook records this.
Selections mode="lab"|"online" and condition="basic"|"enhan" can be
combined with sample. They preserve original IDs/order and filter both
waves together. There are 26 lab cases, all enhanced, and 478 online cases
(221 enhanced, 257 basic); empty selections raise. Among completers there
are 26 lab-enhanced, 155 online-enhanced and 164 online-basic cases. These
counts do not by themselves justify a causal treatment analysis or ignore
selection/dropout. All 504 rows are the source package’s retained sample,
not the entire originally recruited population.
The individual CSV, schema, codebook and export script retain GPL-2.0-or-later
terms. Source numeric values are checked against R doubles, timestamps
against their original instants, and completer frequencies against the
existing aggregate loader. The incomplete-data guide
explains fitting assumptions; cookbook/07_incomplete_probability.py runs
the paired-data and post-only examples offline.
PISA 2003: DAKS response data¶
from knowledgespaces.datasets import load_pisa
from knowledgespaces.derivation import iita
pisa = load_pisa()
assert pisa.data.patterns.shape == (340, 5)
fit = iita(pisa.data, version="minimized")
frequencies = load_pisa(aggregate=True).data
This is the complete five-item response data frame from DAKS 2.1-3:
340 German students in PISA 2003, with the source’s dichotomized scores.
The default preserves every original row and labels a..e. The aggregate
view returns identical-pattern counts, preserving analysis totals.
ResponseDataset exposes data, name, description, source and
source_license. It supplies no ground-truth structure: the example’s
IITA relation is estimated from these responses.
The source does not supply the item wording, original PISA item identifiers,
original polytomous scores or dichotomization rules. The CSV’s source_row
values preserve R row positions, not recovered participant identifiers.
There is no filtering, imputation or new dichotomization in this export.
Do not infer item content or reconstruct scores that are absent from the
source. Loading requires neither R nor internet access.
GPL-2.0-or-later data terms, the source codebook and package metadata are
packaged under knowledgespaces/datasets/data/pisa/, separately from MIT
Python code. All CSV values have been compared with the DAKS source archive;
loader/IITA tests also compare fresh native R counts and all three criteria.
See the frequency and quotient guide for descriptive reports.
Selection from the comparison packages¶
The remaining resources are deliberately outside the bundled data selection. They can motivate later applications without being prerequisites for the algorithms or for the probability case study.
Source and objects |
Decision for 0.2.0 |
Reason |
|---|---|---|
pks: |
Not bundled |
Additional geometry pre/post examples; the chosen probability study already supplies paired records. The geometry export omits its held-out sample. |
pks: |
Not bundled |
Additional 16-item application with corrected source relations. |
pks: |
Not bundled |
Additional science examples whose frequencies were reconstructed from published histograms. |
pks: |
Not bundled |
Additional artificial estimation example; numerical checks already use the textbook and independent synthetic inputs. |
pks: |
Not bundled |
Additional school arithmetic studies. Any later fraction loader must resolve the codebook’s scoring convention explicitly. |
kstMatrix: |
Not bundled |
Additional expert-derived structures and application content, not missing structural APIs. |
kstMatrix: |
Upstream examples acknowledged |
Independently authored examples demonstrate these workflows; their named data objects are not copied. |
kstpy: |
Duplicate component, not bundled |
Its set family equals |
MATLAB: |
Not bundled |
Additional application data and skill maps are unnecessary for the chosen case. No numerical identity with the pks arithmetic data is asserted. |
CbKST: four ODS examples; CDSS: ten CSV/XLSX/ODS examples |
Upstream examples acknowledged |
These teach import, competence mapping and course-skill workflows. Original synthetic fixtures exercise the supported formats and methods without redistributing the course content. |
These are scope decisions, not assertions that omitted resources cannot be redistributed. The pinned pks and DAKS packages declare GPL version 2 or later; kstMatrix, CbKST and CDSS declare GPL version 3. The MATLAB repository declares CC BY-NC-SA 4.0, and no separate dataset-specific grant was established in the inspected metadata. kstpy’s package metadata says LGPL-3.0-or-later, while its README and LICENSE say GPL version 3; that inconsistency remains explicit rather than being resolved by assumption. None of those omitted resources is included merely because a compatible Python method exists.
The release checks compare all included probability and PISA exports against their pinned native sources, validate the packaged checksum manifests, and check the textbook frequencies against pks. Exporting the same source again validates transcription and transformations; it does not establish that the study sample is representative or that a fitted model is scientifically true.