Spreadsheet and table interchange¶
Install the optional engines once:
pip install "knowledgespaces[spreadsheets]"
CSV remains available without these engines. XLSX uses openpyxl and ODS uses odfpy; neither is imported by the pure Python structural and QUERY core. The layouts follow kstIO 0.5-1, CbKST 0.1-1 and CDSS 0.3-1. The package’s frozen tests include files produced by their installed R implementations.
Read and write KST objects¶
from knowledgespaces.io import read_kst, write_kst
space = read_kst("course.xlsx", kind="space", sheet="Knowledge space")
write_kst(space, "course.ods", sheet="Knowledge space")
read_kst requires an explicit interpretation. It checks the corresponding
axioms and rejects nonconforming input. It does not insert endpoint states,
close a relation, replace a family by its union closure or reduce generators
silently. Duplicate response rows retain their multiplicities.
|
Table layout |
Python result |
|---|---|---|
|
One binary column per item |
|
|
Same; empty and full states required |
|
|
Same; union closure additionally required |
|
|
Irredundant nonempty generators covering the domain |
|
|
Square binary matrix, same row/column order |
|
|
Initial item-ID column, binary clause columns |
Original |
|
Same, satisfying the surmise axioms |
|
|
Initial item-ID column, binary skill columns |
Conjunctive |
|
Same, repeated IDs express alternatives |
|
|
Binary respondent or pattern rows |
|
|
Additionally blank cells or a missing marker |
|
Relation entry row a, column b means that a is a prerequisite of b. Reflexivity and transitivity must already hold; equivalences are permitted. Repeated clause IDs are alternatives, not new items. Skill columns can include unused skills. Bases and functions export their compact representations, without enumerating the entire state span.
CSV, XLSX and ODS are selected by extension, or explicitly with format.
sheet accepts an exact name or a zero-based index; CSV uses index 0.
With header=False, pass items to preserve the column labels; otherwise
columns receive numeric string labels. If a header is present, explicit labels
must agree with its order. Labels written as text retain leading zeroes,
whitespace, Unicode and literal = prefixes. Numeric identifiers read from
workbooks are converted to strings; duplicate resulting labels are rejected.
For KST/SRBT/plain binary files, use read_legacy_matrix and its typed
wrappers. A raw legacy matrix can also be checked with
table_to_kst([labels, *matrix], kind="basis") (or another kind above).
Frequencies and missing responses¶
from knowledgespaces.io import read_kst, write_kst
data = read_kst("responses.xlsx", kind="incomplete_data",
count_column="frequency", missing_marker="NA")
write_kst(data, "responses.ods", count_column="frequency")
A frequency column is used only when explicitly named. Exporting an object
with frequencies without count_column raises, preventing accidental loss
of weights or implicit expansion into respondents. Fractional nonnegative
weights are preserved. Without frequencies, every row remains one respondent.
Missing responses are accepted only for incomplete_data. The default text
marker is NA; empty cells are also missing within the data area. Entirely
empty trailing worksheet rows are treated as padding, so exporters use the
explicit marker to retain entirely missing respondents. Never use a missing
marker that also denotes a binary value. No imputation or missingness model
is inferred by importing a file; see incomplete responses.
These missing/frequency extensions are explicit Python layouts, not a claim
that kstIO’s binary kdata reader models incomplete or weighted data.
Multiple objects and curriculum workbooks¶
from knowledgespaces.io import kst_to_table, write_tables
write_tables({"Space": kst_to_table(space),
"Responses": kst_to_table(data, count_column="frequency")},
"study.xlsx")
read_table returns the raw selected table, including any header, for
inspection or editing in Python. table_to_kst performs the subsequent
mathematical validation. write_tables creates a new export workbook; it
does not preserve other sheets, styling or charts from an existing file.
For CDSS pair tables, use read_assignment("course.xlsx").derive().
The default sheet names are Taught and Required, with (object, skill)
columns. taught_sheet and required_sheet select different sheets.
write_assignment exports the same layout. These functions also detect
JSON and two-file CSV inputs. JSON preserves multi-assignments and unused
domain labels that pair tables cannot represent; lossy pair exports raise.
See course-dependent structures and cookbook 11.
Formula results and limits¶
Formulas are rejected by default. formula_policy="cached" explicitly reads
stored results, raising if a cache is absent or erroneous. No formula is
evaluated, and the engines cannot establish whether cached values are current.
This distinction follows the openpyxl reader API.
CSV fields have no formula type and are read as literal text.
max_cells bounds the rectangular area of the selected table, including
interior blanks. Trailing empty formatting padding is ignored; repeated ODS
rows/columns are bounded before expansion. The limit applies per table and
does not bound all memory used internally by a workbook engine. Nonbinary
values are rejected before conversion to integers; dates and error cells are
not silently treated as item responses.
XLSX text is limited to 32,767 characters per cell; oversized text raises instead of being truncated. Workbook exports reject XML-invalid characters. Tabs, line feeds and carriage returns in text cells are preserved.