Patterns, matrices and named sets

Use an explicit item order at the boundaries of a workflow. Ordered response rows and learning paths preserve their duplicates; a SetFamily deliberately deduplicates its members. KnowledgeStructure also ensures the empty/full states are present. These are different operations.

import numpy as np
from knowledgespaces.io import (
    states_to_matrix, matrix_to_states, matrix_to_patterns,
    patterns_to_matrix, expand_response_patterns, subset_matrix,
    fringe_matrix, boolean_matrix_product,
)
from knowledgespaces.estimation import ResponseMatrix
from knowledgespaces.metrics import pattern_frequencies
from knowledgespaces.structures import SetFamily, KnowledgeStructure

items = ["product", "sum", "never observed"]
states = [{"sum"}, {"product", "sum"}, {"sum"}]
rows = states_to_matrix(states, items=items, dtype=bool)
assert rows.tolist() == [[False, True, False], [True, True, False], [False, True, False]]
assert matrix_to_states(rows, items=items) == tuple(map(frozenset, states))
binary = matrix_to_patterns(rows)
assert binary == ("010", "110", "010")
assert np.array_equal(patterns_to_matrix(binary), rows)

text = matrix_to_patterns(rows, items=items, separator=" | ", empty="∅")
assert np.array_equal(patterns_to_matrix(text, items=items, separator=" | ", empty="∅"), rows)
exact = SetFamily(items, matrix_to_states(rows, items=items))
assert len(exact) == 2  # neither endpoint silently inserted

These functions cover the conversion questions of pks::as.pattern/as.binmat and kst::as.binaryMatrix/as.famset. Integer, float and boolean binary inputs are accepted; numeric strings, missing cells and nonbinary scores are not. The output defaults to int8; dtype=bool produces a logical matrix. Empty families retain the supplied column domain. Labelled text requires an unambiguous separator; item names containing the separator and empty symbols indistinguishable from nonempty states are rejected. Binary text uses exactly one bit per item. Naming is explicit: autogenerated R letters and kstpy names are presentation choices, not a mathematical convention of the Python API.

Frequencies and expansion

data = ResponseMatrix(items, rows.astype(np.int8), np.array([2, 1, 3]))
report = pattern_frequencies(data, n=None)
frequencies = dict(zip(matrix_to_patterns(report.patterns), report.counts))
assert frequencies == {"010": 5, "110": 1}
expanded = expand_response_patterns(data, max_rows=10)
assert expanded.shape == (6, 3)

Expansion preserves the original row order, drops zero-count rows and rejects fractional weights. Most estimators accept weighted ResponseMatrix directly, so expansion is optional. max_rows and max_cells guard the allocation rather than returning truncated data. Frequency reports allow fractional analysis weights without converting them to respondent counts.

For projection of response data, select columns in the desired order and retain counts: ResponseMatrix([items[i] for i in selected], data.patterns[:, selected].copy(), data.effective_counts.copy()). If projection creates duplicate rows, these remain valid weighted responses; pattern_frequencies(..., n=None) aggregates them explicitly.

Inclusion, paths and fringes

structure = KnowledgeStructure(["a", "b"], [set(), {"a"}, {"a", "b"}])
ordered = list(structure)
incidence = subset_matrix(ordered)
assert incidence.tolist() == [[True, True, True], [False, True, True], [False, False, True]]
fringes = fringe_matrix(structure, ordered, items=["b", "a"])
assert fringes.tolist() == [[0, 1], [1, 1], [1, 0]]
path = next(structure.iter_learning_paths())
path_matrix = states_to_matrix(path, items=["b", "a"])

subset_matrix has row-to-column inclusion orientation, as in pks::is.subset; proper=True requests strict inclusion. It is not Hasse adjacency. boolean_matrix_product composes binary relations with existential Boolean multiplication and avoids arithmetic-count overflow. fringe_matrix accepts a structure or compact KnowledgeBase without expanding the latter. Its kind can be inner, outer or both; the latter is their union. Row states must belong to the represented structure. These are single-item fringes. The structural reports guide distinguishes paths, multi-item changes and compact neighbourhoods.

For families/bases, relations, clauses, assignments and typed response tables, continue to use kst_to_table/table_to_kst; these also provide readable named rows. No second table or spreadsheet system is introduced. Use SurmiseRelation.to_matrix() for a named reflexive relation matrix; the orientation is prerequisite row to dependent column.

Legacy KST compatibility

read_legacy_matrix(..., kind="data", format="KST", items=items) and write_legacy_matrix(items, rows, ..., kind="data", format="KST") retain row order and duplicates in the legacy item-count/row-count/binary-row format used by kstpy. It stores no item labels or weights. Expand integer frequencies explicitly before writing, or use a named table/JSON to retain weights. read_legacy_structure requires endpoint states; use the matrix reader and SetFamily for arbitrary families. Default labels differ across software: supply the intended mapping rather than guessing it from a filename.

Files with ambiguous headerless layouts require an explicit format. Source checks and native round trips are restricted to the archived versions used for the release comparison; exact console printing is not promised.

In kstpy 1.0.0 the native set writer produces readable KST files, but its list writer and both readers fail internal helper type checks on their documented inputs. Our reader was verified on the native set writer; bidirectional native execution is not claimed. The Python round trip retains duplicates. Similarly, kstMatrix 3.0-0’s Boolean-product helper gives the wrong output dimensions for rectangular matrices; Python uses the standard outer dimensions and the existential composition definition.

The runnable cookbook/18_family_interoperability.py covers exact observed families, minimal/maximal structures, table conversions, valid implications on arbitrary families and projection across object types. Empty-domain projection remains an exact SetFamily; KST estimators require a nonempty domain.