Random starts and reproducible refits

BLIM likelihoods can have local optima. Random starts explore several initializations; they provide no guarantee of a global maximum or identified parameters. Keep the data, constraints, initialization law, iteration cap and tolerance with the result.

Initialization laws

estimate_blim_restarts and estimate_blim_incomplete_restarts accept two strategies:

Strategy

Error parameters

State prior

uniform (default)

Draw beta and eta independently from init_range, default (0.01, 0.4); halve both at an item while their sum is at least .95.

Equal mass on each state.

pks

Draw beta and eta independently from U(0,1); reflect beta when their sum is at least 1, then reflect eta if the updated sum is still at least 1.

Dirichlet(1, …, 1), uniform on the simplex.

The second prior has the same distribution as the spacings of sorted uniform points used in pks 0.7-0 blim(randinit=TRUE). This is a law of starting values, not an equivalence of R and NumPy random streams. In particular, normalizing independent U(0,1) values would give a different prior law. Fixed state probabilities override this draw. Fixed/shared error constraints project the error starts; free errors use the estimator’s numerical box. The informative-start inequality does not constrain subsequent EM iterates.

For SLM, estimate_slm_restarts draws solvabilities from U(0,1); its state prior follows from these solvabilities, not from a free simplex draw. The deterministic MD estimator needs no random restarts.

Bootstrap policy

bootstrap_blim_incomplete(..., n_restarts=None) retains the default single deterministic start for both the observed data and each simulated sample. Setting n_restarts to a positive integer uses that many random starts for every fit, including the observed fit. The same init_strategy, init_range, constraints, tolerance and cap apply throughout.

from knowledgespaces.estimation import bootstrap_blim_incomplete

boot = bootstrap_blim_incomplete(
    structure, data, constraints=constraints, mask_model="fixed",
    n_replicates=4, n_restarts=3, init_strategy="pks", seed=63,
)
assert boot.n_failed == boot.n_capped == 0
assert len(boot.estimate.restarts) == 3
assert all(len(records) == 3 for records in boot.restart_replicates)

The seeded generator is shared by the observed search, simulations and refits in that order. Consequently, changing the search policy also changes later random simulations. Neither function changes NumPy’s global random state. restart_replicates retains each replicate’s search; replicate_errors distinguishes samples without observed responses from all-start failures. Failed rows remain NaN, capped fits remain visible, and the bootstrap p-value is unavailable when an observed or selected replicate fit is unresolved. Unselected unsuccessful starts do not invalidate a successfully selected fit.

The explicit observation-mask model remains essential. A fixed mask is a fixed-design/MCAR simulation; a callback defines its own missingness model. Multistart optimization does not establish ignorability or identify an MNAR mechanism. See the incomplete-data guide.

Verification scope

Six native pks 0.7-0 random-start BLIM fits (ML/MDML, three seeds) are compared after exporting their actual initial parameters; 17 EM iterations agree numerically. The tests separately check the uniform-simplex moments, incomplete-fit replay, analytic one-item likelihood, constraints, numerical failures and exact bootstrap refit policies. These are finite numerical checks, not proofs that EM always converges or reaches the global optimum.