Skip to content
Merged
115 changes: 115 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# Changelog

## 0.3.0 (unreleased)

A correctness review and breaking refactor. Old names are **not** kept as
aliases; use the tables below to migrate.

### Behavior changes

- **All information quantities are in bits**, including log-likelihoods
(`log_likelihood`, the Baum-Welch trace, cross-validated log-likelihood,
HDP-HMM traces), topological entropy (was nats), and collision entropy (was
nats). AIC, AICc, BIC, and WAIC keep their standard deviance scale (computed
from the natural log-likelihood); the MDL code length is in bits. `score` and
`observed_information` remain derivatives of the natural log-likelihood, so
standard errors are unchanged.
- **Covers follow Lind & Marcus**: *right* means right-resolving.
`RightFischerCover` is now the exact minimal right-resolving presentation
(it was built from follower sets truncated at length 8 and labeled *left*);
reducible shifts raise `SoficValidationError`. `LeftFischerCover` is its
mirror image.
- `MaximizedPrimeAtomaton` subclasses `NFA`, not `AtomicAutomaton`: it is the
reverse of the canonical RFSA of the reversed language and need not be atomic.
- `CanonicalVisiblyPushdownAutomaton.from_vpa` raises
`NonWellMatchedLanguageError` for languages with pending calls or returns.
- `DeterministicVisiblyPushdownAutomaton.from_vpa` determinizes
nondeterministic input instead of raising.
- VPA boolean and structural operations return concrete automata.
- Wildcard VPA returns fire on every stack symbol and, with a bottom symbol, on
the empty stack (the simulator's behavior; the docstring said otherwise).
- `ShiftOfFiniteType.from_forbidden_words` raises instead of silently
truncating at `max_states`; an empty forbidden set builds the full shift.
- `equivalent()` compares automata over the union of both alphabets.
- `MarkovChain.words_of_length(0)` and `sample_path` respect the initial
distribution; sampling from an initial law with no mass raises `ValueError`.
- `block_entropy_estimates(use_exact=True)` reports crypticity as `C_mu - E`.
- `SlidingBlockCode.apply` keeps constraints longer than the window and raises
when the block map misses an allowed block.
- Baum-Welch raises when every sequence is impossible and warns when some are.
- `BuchiAutomaton.accepts_lasso` rejects an empty loop.
- Stack CSSR no longer drops histories during determinization (which produced
zero-mass states and validation errors).
- TikZ labels: fixed uncompilable `\midcall` / `\uparrowA` and escaped
`\times` / sympy LaTeX.
- cmpy example factories validate parameters like their curated counterparts
(e.g. a coin bias of 1 raises). Processes previously built by `Even`, `Nemo`,
and `ABC` now come from the curated functions and emit integer symbols `0`/`1`;
`alternating_biased_coins(p, p)` keeps its two-state presentation.
- `joint_block_distribution(block_length=n)` counts blocks of `n` symbols
(`history_length=h` meant `h + 1` symbols); the default `2` is unchanged.

### New

- Exact canonical RFSA (`CanonicalRFSA.from_language`), maximized prime
átomaton, `CanonicalRFSA.dual` / `MaximizedPrimeAtomaton.dual`, exact
`prime_residuals`, `atoms`, `prime_atoms`, and `ResidualTable`.
- NL\*: `learn_rfsa_nlstar`, `learn_prime_atomaton_nlstar`,
`learn_rfsa_from_language`, and `AutomatonEquivalenceOracle`.
- VPA: `determinize`, `is_empty`, `accepted_word`, `is_universal`, `includes`,
`equivalent`, `has_unmatched_word`; `sofic.automata.vpa.to_single_entry` /
`to_multiple_entry`; modular `minimize` converts automatically when no
modules are given. `NestedWordAutomaton` gains the same operations.
- Exact left/right Krieger covers.
- `sofic.generators.matrices` (joint matrices and start-vector policies) and
`sofic.generators.sampling`.

### Module moves

| Old module | New module |
|---|---|
| `sofic.generators.hmm_inference` | `sofic.inference.hmm` (`filtering`, `em`, `information`); `sample` → `sofic.generators.sampling` |
| `sofic.generators.epsilon_inference` | `sofic.inference.cssr` (`process`, `subtree`, `counts`, `significance`); `spectral` → `sofic.inference.spectral` |
| `sofic.generators.epsilon_transducer_inference` | `sofic.inference.cssr.transducer` |
| `sofic.generators.stack_inference` | `sofic.inference.cssr.stack` |
| `sofic.automata.{active,rpni,edsm,dfasat,alergia,papni,observation}` | `sofic.automata.learning.*` |
| `sofic.automata.learning` (NL\*) | `sofic.automata.learning.nlstar` |
| `sofic.automata.vpa`, `vpa_simulation` | `sofic.automata.vpa` package (`base`, `operations`, `deterministic`, `modular`, `canonical`, `simulation`) |
| `sofic.automata.{icdfa,idfa,enumeration}` | `sofic.automata.enumeration.{icdfa,idfa,words}` |
| `sofic.automata.{canonical_extraction,rfsa,atomaton,canonical_dual}` | `sofic.automata.canonical.{residual,rfsa,atomaton,dual}` |
| `sofic.shifts.sofic_relation` | `sofic.shifts.product_alphabet_shift` |

### Renames and removals

| Old | New |
|---|---|
| `cssr` | `learn_epsilon_machine_cssr` |
| `subtree_merge` | `learn_epsilon_machine_subtree` |
| `spectral` (wrapper) | `learn_epsilon_machine_spectral` |
| `transcssr` | `learn_epsilon_transducer_cssr` |
| `stack_cssr` / `stack_subtree_merge` / `fit_stack_hmm_mle` | `learn_stack_hmm_cssr` / `learn_stack_hmm_subtree` / `learn_stack_hmm_mle` |
| `suggest_lmax` | `suggest_max_history` |
| `Lmax=`, `L=` (CSSR family and subtree learners) | `max_history=` |
| `reconstruction_sweep(lmaxes=)` | `reconstruction_sweep(max_histories=)` |
| `GoodnessOfFit.L`, `goodness_of_fit(L=)` | `block_length` |
| `forward(scaled=)`, `backward(scaled=)` | `normalize=` |
| `joint_block_distribution(history_length=h)` | `joint_block_distribution(block_length=h + 1)` |
| `QuasiStochasticModel.transition_matrices`, `quasi_inference.transition_matrices` | `symbol_matrices` |
| `sofic.generators.words.hmm_*`, `pfa_*`, `quasi_*`, `markov_*` functions | private; use the model methods |
| `cartesian_product_gg` / `cartesian_product_tt` | `generator_product` / `transducer_product` |
| `compose_tt` / `compose_tg` | `compose_transducers` / `compose_transducer_generator` |
| `wnfa_to_wdfa` / `minimum_wdfa` | `determinize_wheeler` / `minimize_wheeler` |
| `papni_encode` / `papni_encode_samples` | `encode_dyck_word` / `encode_dyck_samples` |
| ICDFA helpers `next_flags`, `string_from_flags`, `flags_from_string`, `count_flag_sequences` | `icdfa_next_flags`, `icdfa_string_from_flags`, `icdfa_flags_from_string`, `icdfa_count_flag_sequences` |
| IDFA helpers `string_from_flags`, `extended_flags`, `transition_count` | `idfa_string_from_flags`, `idfa_extended_flags`, `idfa_transition_count` |
| `CallDrivenAutomaton` | `ModularVisiblyPushdownAutomaton` |
| `CompositeVisiblyPushdownAutomaton`, `union_vpa`, `intersection_vpa`, `complement_vpa`, `difference_vpa`, `concat_vpa`, `kleene_star_vpa` | removed; use the VPA methods |
| `LabeledAutomaton.intersect` / `concatenate` / `star` | `intersection` / `concat` / `kleene_star` |
| `learn_maximized_prime_atomaton` | `learn_prime_atomaton_nlstar` (or `learn_rfsa_nlstar`) |
| `SoficRelation`, `to_sofic_relation` | `ProductAlphabetShift`, `to_product_alphabet_shift` |
| cover `from_sofic` | `from_presentation` |
| `{left,right}_{fischer,krieger}_from_sofic` | `{left,right}_{fischer,krieger}_cover` |
| `sofic.from_yaml`, `serialization.from_yaml` | `model_from_yaml` |
| `examples.processes`: `BiasedCoin`, `FairCoin`, `Even`, `Nemo`, `NRPS`, `ABC` | removed: `bernoulli`, `fair_coin`, `even_process`, `nemo_process`, `noisy_random_phase_slip`, `alternating_biased_coins(1 - p, 1 - q)` |
| `GoldenMean` / `Butterfly` / `PSB` | `golden_mean_forbid_00` / `butterfly_two_branch` / `phase_slip_backtrack_cmpy` |
| other PascalCase `examples.processes` factories (e.g. `BitFlip`, `Delay`, `RestrictedGM`, `IrreversibleTwoState`) | snake_case (`bit_flip`, `delay`, `restricted_gm`, `irreversible_two_state`) |
2 changes: 1 addition & 1 deletion README.rst
Original file line number Diff line number Diff line change
Expand Up @@ -171,7 +171,7 @@ read off computational-mechanics quantities:
from sofic import EpsilonMachine

eps = EpsilonMachine.from_hmm(gm) # minimize an HMM presentation
# eps = EpsilonMachine.from_sequence(data, method="cssr", Lmax=4) # infer
# eps = EpsilonMachine.from_sequence(data, method="cssr", max_history=4) # infer
# eps = EpsilonMachine.from_sequence(data, method="spectral", prefix_length=3, rank=2)

eps.statistical_complexity() # 0.9183 bits (C_mu)
Expand Down
18 changes: 16 additions & 2 deletions docs/automata/atomaton.rst
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
.. atomaton.rst
.. py:module:: sofic.automata.atomaton
.. py:module:: sofic.automata.canonical.atomaton

********
Átomaton
Expand Down Expand Up @@ -29,13 +29,24 @@ the special case where the reverse is deterministic.

.. code-block:: python

from sofic.automata.atomaton import Atomaton, atomic_states, is_atomic
from sofic.automata.canonical.atomaton import Atomaton, atomic_states, is_atomic

atomaton = Atomaton.from_language(dfa)
is_atomic(atomaton) # True
atomic_states(nfa) # states whose right language is a union of atoms
is_atomic(nfa.reverse()) # iff nfa.determinize() is minimal

Maximized prime átomaton
========================

The maximized prime átomaton (:class:`MaximizedPrimeAtomaton`) is the dual of
the canonical RFSA :cite:`MaarandTamm2022`: the reverse of the canonical RFSA of
the reversed language, just as the átomaton is the reverse of the minimal DFA of
the reversed language. Its states are the maximized prime atoms, and the right
language of each lies between its atom and its maximized atom :cite:`Tamm2015`.
Unlike the átomaton it need not be atomic, so it is a plain
:class:`~sofic.automata.nfa.NFA` subclass.

API
===

Expand All @@ -45,3 +56,6 @@ API
.. autoclass:: AtomicAutomaton
.. autoclass:: Atomaton
.. autoclass:: MaximizedPrimeAtomaton
:members: from_language, from_canonical_rfsa, dual

.. autofunction:: sofic.automata.canonical.residual.maximized_prime_atomaton_from_language
2 changes: 1 addition & 1 deletion docs/automata/dfa.rst
Original file line number Diff line number Diff line change
Expand Up @@ -39,4 +39,4 @@ API
===

.. autoclass:: DFA
:members: add_transition, recognizes, union, intersection, intersect, complement, difference, concat, concatenate, kleene_star, star, minimize, from_nfa
:members: add_transition, recognizes, union, intersection, complement, difference, concat, kleene_star, minimize, from_nfa
10 changes: 5 additions & 5 deletions docs/automata/icdfa.rst
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
.. icdfa.rst
.. py:module:: sofic.automata.icdfa
.. py:module:: sofic.automata.enumeration.icdfa

*****
ICDFA
Expand All @@ -23,7 +23,7 @@ API
.. autofunction:: count_icdfa
.. autofunction:: count_icdfa_empty

.. autofunction:: sofic.automata.idfa.iter_idfa_strings
.. autofunction:: sofic.automata.idfa.rank_idfa_string
.. autofunction:: sofic.automata.idfa.unrank_idfa_string
.. autofunction:: sofic.automata.idfa.count_accessible_idfa
.. autofunction:: sofic.automata.enumeration.idfa.iter_idfa_strings
.. autofunction:: sofic.automata.enumeration.idfa.rank_idfa_string
.. autofunction:: sofic.automata.enumeration.idfa.unrank_idfa_string
.. autofunction:: sofic.automata.enumeration.idfa.count_accessible_idfa
81 changes: 46 additions & 35 deletions docs/automata/learning.rst
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
.. learning.rst
.. py:module:: sofic.automata.learning
.. py:module:: sofic.automata.learning.nlstar

********
Learning
Expand All @@ -10,23 +10,33 @@ Learning
Active learning (NL\*)
======================

Active learning of maximized prime átomatons via NL\* with a membership
teacher, following Angluin-style learning and its nondeterministic extension
:cite:`Angluin1987,Bollig2009`:
NL\* :cite:`Bollig2009` extends Angluin's L\* :cite:`Angluin1987` to
nondeterministic automata. It keeps an RFSA-closed, RFSA-consistent observation
table whose prime rows become the hypothesis states, and adds every suffix of a
counterexample as a new experiment. When the equivalence oracle accepts, the
hypothesis is the canonical RFSA of the target (:doc:`rfsa`). Running NL\* on
the reversed target and reversing the result learns the maximized prime
átomaton (:doc:`atomaton`).

.. autofunction:: sofic.automata.learning.learn_maximized_prime_atomaton
:class:`~sofic.automata.learning.active.AutomatonEquivalenceOracle` answers equivalence
queries exactly against a target automaton, returning a shortest
counterexample.

.. autofunction:: sofic.automata.learning.nlstar.learn_rfsa_nlstar
.. autofunction:: sofic.automata.learning.nlstar.learn_prime_atomaton_nlstar
.. autofunction:: sofic.automata.learning.nlstar.learn_rfsa_from_language

Active learning (L\*, TTT, Mealy)
=================================

:mod:`sofic.automata.active` learns an automaton from a *teacher* answering
:mod:`sofic.automata.learning.active` learns an automaton from a *teacher* answering
**membership** and **equivalence** queries. It provides Angluin's L\*
:cite:`Angluin1987` and a redundancy-free **discrimination-tree** learner in the
TTT family :cite:`KearnsVazirani1994,Isberner2014` for
:class:`~sofic.automata.dfa.DFA`, plus the Mealy variant of L\*
:cite:`Shahbaz2009`; all use Rivest-Schapire counterexample analysis
:cite:`RivestSchapire1993`. Oracles adapt sofic models: a
:class:`~sofic.automata.active.LanguageMembershipOracle` wraps any model exposing
:class:`~sofic.automata.learning.active.LanguageMembershipOracle` wraps any model exposing
``recognizes`` / ``__contains__`` (a :class:`~sofic.automata.dfa.DFA`, NFA,
átomaton, or ``model.to_support_dfa()`` for a sofic shift or ε-machine), and the
equivalence oracles offer bounded-exhaustive or random-walk testing.
Expand All @@ -43,7 +53,7 @@ equivalence test:

.. code-block:: python

from sofic.automata.active import (
from sofic.automata.learning.active import (
FunctionMembershipOracle,
RandomWalkEquivalenceOracle,
learn_dfa_lstar,
Expand All @@ -53,22 +63,23 @@ equivalence test:
equivalence = RandomWalkEquivalenceOracle(membership, {"a", "b"}, rng=0)
dfa = learn_dfa_lstar({"a", "b"}, membership, equivalence)

.. autofunction:: sofic.automata.active.learn_dfa_lstar
.. autofunction:: sofic.automata.active.learn_dfa_ttt
.. autofunction:: sofic.automata.active.learn_mealy_lstar
.. autofunction:: sofic.automata.active.learn_dfa_from_language
.. autofunction:: sofic.automata.active.learn_mealy_from_transducer
.. autofunction:: sofic.automata.learning.active.learn_dfa_lstar
.. autofunction:: sofic.automata.learning.active.learn_dfa_ttt
.. autofunction:: sofic.automata.learning.active.learn_mealy_lstar
.. autofunction:: sofic.automata.learning.active.learn_dfa_from_language
.. autofunction:: sofic.automata.learning.active.learn_mealy_from_transducer

.. autoclass:: sofic.automata.active.MembershipOracle
.. autoclass:: sofic.automata.learning.active.MembershipOracle
:members:
.. autoclass:: sofic.automata.active.EquivalenceOracle
.. autoclass:: sofic.automata.learning.active.EquivalenceOracle
:members:
.. autoclass:: sofic.automata.active.LanguageMembershipOracle
.. autoclass:: sofic.automata.active.FunctionMembershipOracle
.. autoclass:: sofic.automata.active.ExhaustiveEquivalenceOracle
.. autoclass:: sofic.automata.active.RandomWalkEquivalenceOracle
.. autoclass:: sofic.automata.active.TransducerOutputOracle
.. autoclass:: sofic.automata.active.MealyExhaustiveEquivalenceOracle
.. autoclass:: sofic.automata.learning.active.LanguageMembershipOracle
.. autoclass:: sofic.automata.learning.active.FunctionMembershipOracle
.. autoclass:: sofic.automata.learning.active.AutomatonEquivalenceOracle
.. autoclass:: sofic.automata.learning.active.ExhaustiveEquivalenceOracle
.. autoclass:: sofic.automata.learning.active.RandomWalkEquivalenceOracle
.. autoclass:: sofic.automata.learning.active.TransducerOutputOracle
.. autoclass:: sofic.automata.learning.active.MealyExhaustiveEquivalenceOracle

Passive learning (RPNI)
=======================
Expand All @@ -84,7 +95,7 @@ consistent with the sample :cite:`Lang1998`:
dfa = learn_dfa_rpni(positive=["ab", "abab"], negative=["a", "b"])
dfa.validate()

.. autofunction:: sofic.automata.rpni.learn_dfa_rpni
.. autofunction:: sofic.automata.learning.rpni.learn_dfa_rpni

Passive learning (EDSM / blue-fringe)
=====================================
Expand All @@ -105,12 +116,12 @@ automaton:
dfa = learn_dfa_edsm(positive=["a", "aba", "ababa"], negative=["", "b", "ab"])
dfa.validate()

.. autofunction:: sofic.automata.edsm.learn_dfa_edsm
.. autofunction:: sofic.automata.learning.edsm.learn_dfa_edsm

Exact minimal DFA (SAT)
=======================

Where RPNI and EDSM are heuristics, :func:`sofic.automata.dfasat.learn_dfa_sat`
Where RPNI and EDSM are heuristics, :func:`sofic.automata.learning.dfasat.learn_dfa_sat`
returns the **provably minimal** DFA consistent with the sample. Following Heule
& Verwer :cite:`HeuleVerwer2010`, it translates the augmented prefix-tree
acceptor into a graph-colouring SAT instance and searches the state count ``k``
Expand All @@ -125,7 +136,7 @@ satisfiable ``k``. It requires the optional `python-sat
dfa = learn_dfa_sat(positive=["a", "aba", "ababa"], negative=["", "b", "ab"])
dfa.validate()

.. autofunction:: sofic.automata.dfasat.learn_dfa_sat
.. autofunction:: sofic.automata.learning.dfasat.learn_dfa_sat

Probabilistic passive learning (ALERGIA)
========================================
Expand All @@ -135,7 +146,7 @@ from **unlabeled** positive strings by merging states of a frequency
prefix-tree acceptor whenever a Hoeffding-bound test cannot distinguish their
transition statistics :cite:`Carrasco1994`. It is the stochastic, unlabeled
counterpart of RPNI/EDSM and a state-merging alternative to CSSR
(:func:`sofic.generators.epsilon_inference.cssr`). The compatibility threshold
(:func:`sofic.inference.cssr.process.cssr`). The compatibility threshold
``alpha`` trades off model size against fidelity: smaller ``alpha`` merges more
aggressively (fewer states); larger ``alpha`` is more conservative.

Expand All @@ -148,13 +159,13 @@ aggressively (fewer states); larger ``alpha`` is more conservative.
pfa = learn_pfa_alergia(samples, alpha=0.05)
pfa.validate()

.. autofunction:: sofic.automata.alergia.learn_pfa_alergia
.. autofunction:: sofic.automata.learning.alergia.learn_pfa_alergia

Passive learning (PAPNI)
========================

PAPNI extends passive inference to visibly pushdown languages. Words over a
:class:`~sofic.automata.papni.DyckAlphabet` are stack-encoded, a DFA is
:class:`~sofic.automata.learning.papni.DyckAlphabet` are stack-encoded, a DFA is
induced over the encoding, and the result is decoded to a
:class:`~sofic.shifts.sofic_dyck.SoficDyckShift` :cite:`Muskardin2025`:

Expand All @@ -170,13 +181,13 @@ induced over the encoding, and the result is decoded to a
shift = learn_sofic_dyck_shift_papni(positive=["()", "(())"], negative=["("], alphabet=alphabet)

For fitting probabilities on the learned topology, see
:doc:`../generators/stack_inference`.
:doc:`../inference/stack_cssr`.

.. autoclass:: sofic.automata.papni.DyckAlphabet
.. autoclass:: sofic.automata.learning.papni.DyckAlphabet
:members: classify, symbol_alphabet

.. autofunction:: sofic.automata.papni.learn_sofic_dyck_shift_papni
.. autofunction:: sofic.automata.papni.is_well_matched
.. autofunction:: sofic.automata.papni.papni_encode
.. autofunction:: sofic.automata.papni.papni_encode_samples
.. autofunction:: sofic.automata.papni.sofic_dyck_shift_from_papni_dfa
.. autofunction:: sofic.automata.learning.papni.learn_sofic_dyck_shift_papni
.. autofunction:: sofic.automata.learning.papni.is_well_matched
.. autofunction:: sofic.automata.learning.papni.encode_dyck_word
.. autofunction:: sofic.automata.learning.papni.encode_dyck_samples
.. autofunction:: sofic.automata.learning.papni.sofic_dyck_shift_from_papni_dfa
2 changes: 1 addition & 1 deletion docs/automata/nfa.rst
Original file line number Diff line number Diff line change
Expand Up @@ -31,4 +31,4 @@ API
===

.. autoclass:: NFA
:members: add_transition, recognizes, union, intersection, intersect, complement, difference, concat, concatenate, kleene_star, star, determinize, minimize
:members: add_transition, recognizes, union, intersection, complement, difference, concat, kleene_star, determinize, minimize
Loading
Loading