Visualization

The knowledgespaces.viz module draws structures, relations, clauses and statistical reports with matplotlib. It is part of the optional viz extra:

pip install "knowledgespaces[viz]"

Importing knowledgespaces.viz does not require matplotlib. Calling a plotting function without it raises an ImportError with the install hint. Numerical exports such as hasse_data work without the plotting dependency.

Hasse diagram of a structure

plot_hasse lays members out by size and draws inclusion covers: an edge means no other plotted member lies strictly between its endpoints. A cover can add several items in a general structure. It accepts KnowledgeStructure, SetFamily and KnowledgeBase; a base plots its rows without expanding the space, and an arbitrary family gains no extra states.

import knowledgespaces as ks
from knowledgespaces.viz import plot_hasse

structure = ks.space_from_prerequisites(
    ["add", "sub", "mul"], [("add", "sub"), ("sub", "mul")])
fig = plot_hasse(structure)
fig.savefig("hasse.png", dpi=200)

Useful keyword arguments:

  • highlight_states — a collection of states to color differently (for example the base, the atoms at an item, or the state returned by an assessment);

  • figsize and font_size — control the physical size and label size; for print, prefer a small figsize (for example (5.5, 3.5)) with font_size=9, so that labels remain readable after page scaling;

  • title — None (default) draws a standard title, a custom string replaces it, and the empty string "" suppresses the title;

  • ax — draw into an existing matplotlib Axes for composite figures.

  • orientation="horizontal" — put increasing size from left to right;

  • node_labels and node_colors — mappings keyed by exact states, preserving their identities independently of display names;

  • node_values — one finite value per state, mapped to a colormap and labelled colorbar (cmap, value_label). Suitable for state masses, but values are not automatically normalized or interpreted as probabilities;

  • highlight_paths — collections of state sequences, with their cover edges highlighted; invalid members or skipped covers are rejected;

  • centre — mark a member with a diamond, useful for neighbourhoods.

Node colors or values remain visible when a state is highlighted: an outline marks it. With no custom colors, highlight fill also changes. node_colors and node_values are mutually exclusive. Labels, title and layout are presentation choices; they never change the family.

Paths, neighbourhoods and numerical graph export

import json
from knowledgespaces import KnowledgeBase, KnowledgeStructure, SetFamily
from knowledgespaces.viz import hasse_data, plot_hasse

structure = KnowledgeStructure("abc", ["a", "ab", "ac"])
path = next(structure.iter_learning_paths(through="ab"))
path_fig = plot_hasse(structure, highlight_paths=[path], orientation="horizontal")

centre = frozenset("a")
local = SetFamily(structure.domain, structure.neighbourhood(centre, include=True))
colors = {
    state: "#E74C3C" if state == centre else ("#2ECC71" if centre < state else "#4A90D9")
    for state in local
}
local_fig = plot_hasse(local, centre=centre, node_colors=colors, title="Neighbourhood of {a}")

graph = hasse_data(local).as_dict()
graph_json = json.dumps(graph)  # named states, cover-edge indices, coordinates and domain
assert len(graph["states"]) == len(local)

The local diagram represents covers within the selected family; it is not an automatic assertion that every displayed edge is a global cover in a larger structure. Path lists/matrices can be plotted by explicitly forming a SetFamily from their states, or highlighted in the full structure. The structural report guide covers path filtering and matrix conversion. A large state space or factorial path enumeration is not made cheap by plotting it; select members explicitly when needed.

Hasse diagram of a surmise relation

plot_relation draws the covering relation (Hasse diagram) of a SurmiseRelation:

from knowledgespaces.structures import SurmiseRelation
from knowledgespaces.viz import plot_relation

rel = SurmiseRelation(
    ["add", "sub", "mul"], [("add", "sub"), ("sub", "mul")])
fig = plot_relation(rel)

Both functions return the matplotlib Figure, so they compose with the usual matplotlib workflow (savefig, subplots, style contexts). Vector output (.pdf, .svg) is recommended for publication figures.

plot_relation supports the same two orientations and item-keyed node_labels/node_colors. With cycles, use collapse_equivalent=True to plot the partial-order quotient; mappings then use its representative item labels (relation.quotient()). Arrows always point from prerequisite to successor, independent of orientation.

Clauses and atomic items

from knowledgespaces.viz import plot_surmise_function

function = KnowledgeBase("abc", ["ab", "c"]).surmise_function()
clause_fig = plot_surmise_function(function)
space_fig = plot_surmise_function(function, space=True, max_items=4)

The default plots the base. Each clause is labelled with every item at which it is an atom; {a,b} above is atomic at both a and b. space=True explicitly expands the span, subject to the usual item guard (max_items is exclusive). The clause families themselves remain available through clauses_for(item) for numerical export. This covers the semantic content of kstMatrix’s surmise-function plot without reproducing its HTML/DOT label syntax.

Coefficients and bootstrap distributions

The plotting functions take explicit named coefficients and numerical arrays, so they work with complete/incomplete BLIM, SLM and model-comparison results without fitting anything again. Array columns must match the mapping’s insertion order. The following small example is illustrative, not empirical data:

import numpy as np
from knowledgespaces.viz import plot_coefficients, plot_bootstrap_distribution

coefficients = {"a": 0.1, "b": 0.2}
replicates = np.array([[0.08, 0.18], [0.1, 0.2], [0.14, 0.25], [np.nan, np.nan]])
converged = np.array([True, True, False, False])
coef_fig = plot_coefficients(coefficients, replicates=replicates,
                             converged=converged, kind="box", title="Careless errors")
sd_fig = plot_coefficients(coefficients, standard_deviations={"a": 0.03, "b": 0.04})

statistics = np.array([0.5, 1.0, 1.0, 20.0, np.nan])
stat_fig = plot_bootstrap_distribution(statistics, observed=2.0, xlabel="G²")
tail_fig = plot_bootstrap_distribution(statistics, observed=2.0, kind="survival", xlabel="G²")

For actual fit results, construct the coefficient mapping from fit.beta_dict(), fit.eta_dict(), or slm_fit.g_dict(). Pass the corresponding bootstrap.beta_replicates, eta_replicates or g_replicates; rows preserve the original replication order. Complete-data bootstrap parameter_summary() provides empirical SDs with explicit selection policy. Its SDs describe resampling variability, not automatic confidence intervals or evidence of identifiability. SD bars are ±1 SD and are not clipped to [0,1].

For a complete-data GOF result, pass its g2_replicates, g2_observed and converged_replicates. An incomplete-data result’s observed value is observed_gof.G2 for the G2 procedure. For model comparisons, use statistics, statistic and the conjunction of the two fit-convergence flags per replicate. Pearson statistics can be plotted in the same way when provided by the fitted procedure. These functions do not recompute p-values.

Boxplots and distribution plots display all finite replicate values, including nonconverged refits and numeric outliers. Point plots use supplied coefficients and optional SDs; a replicate matrix supplied to a point plot adds status information only, without plotting its distribution or computing SDs automatically. Nonfinite values cannot be plotted; their count and the supplied nonconvergence count remain visible. Boxplots keep outliers. No input array is modified. The survival curve uses the fraction of finite values greater than or equal to each unique value, preserving ties; if failures are present, this is explicitly a distribution conditional on finite values. It is not a valid repaired bootstrap p-value and is not the add-one Monte Carlo calculation. A wholly nonfinite sample produces a labelled empty figure. Undefined SDs are likewise annotated rather than replaced by zero.

The original result arrays remain the numerical record. For example, an export containing every parameter row, including NaNs and convergence flags:

import io

export = np.column_stack([replicates, converged.astype(int)])
buffer = io.StringIO()
np.savetxt(buffer, export, delimiter=",", header="beta_a,beta_b,converged", comments="")
assert "nan" in buffer.getvalue()

This covers the error-bar/box/distribution capabilities of the pinned MATLAB plotblim source. Python deliberately does not copy its outlier-removal logic or imply native MATLAB execution. Layout, palette, file export and cosmetic conventions are delegated to matplotlib. Existing plot_blim_residuals complements these figures with signed Pearson/deviance residuals against fitted probabilities.