Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/automata/dfa.rst
Original file line number Diff line number Diff line change
Expand Up @@ -39,4 +39,4 @@ API
===

.. autoclass:: DFA
:members: add_transition, recognizes, union, intersection, intersect, complement, difference, concat, concatenate, kleene_star, star, minimize, from_nfa
:members: add_transition, recognizes, union, intersection, complement, difference, concat, kleene_star, minimize, from_nfa
2 changes: 1 addition & 1 deletion docs/automata/nfa.rst
Original file line number Diff line number Diff line change
Expand Up @@ -31,4 +31,4 @@ API
===

.. autoclass:: NFA
:members: add_transition, recognizes, union, intersection, intersect, complement, difference, concat, concatenate, kleene_star, star, determinize, minimize
:members: add_transition, recognizes, union, intersection, complement, difference, concat, kleene_star, determinize, minimize
13 changes: 13 additions & 0 deletions docs/examples.rst
Original file line number Diff line number Diff line change
Expand Up @@ -181,6 +181,19 @@ but accepting a ``machine_type`` argument:
In [3]: gm.entropy_rate()
Out[3]: 0.6666666666666665

Several factories share one implementation with a curated example and differ only
in string symbols (``"0"``/``"1"``), state names, or parametrization:
``BiasedCoin(b)`` is :func:`bernoulli` ``(b)``, ``Even`` is :func:`even_process`,
``Nemo`` is :func:`nemo_process`, ``GoldenMean(b)`` is
:func:`golden_mean_forward` ``(1 - b)`` (and :func:`golden_mean_markov` ``(b)``
with ``A``/``B`` swapped), ``ABC(p, q)`` is :func:`alternating_biased_coins`
``(1 - p, 1 - q)`` for ``p != q``, ``RestrictedGM`` is
:func:`restricted_golden_mean`, and ``IrreversibleTwoState()`` is
:func:`ellison_fig9_forward`. Similar names do not always mean the same process:
:func:`golden_mean` is the ``0 <-> 1`` mirror of ``GoldenMean``, and ``Butterfly``
and ``PSB`` differ from :func:`butterfly_process` and
:func:`~sofic.examples.epsilon_machines.phase_slip_backtrack`.

The module also exposes registries — ``processes.process_list`` and
``processes.transducer_list`` — that enumerate every factory, which is handy for
parametrized tests and sweeps:
Expand Down
6 changes: 5 additions & 1 deletion docs/generators/epsilon_inference.rst
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,11 @@ it as the longest history the data supports, not as a synchronization length.
The morph tests also rely on the chi-squared limit, which fails for the sparse
counts of long suffixes. ``test="exact"`` compares the G statistic with tables
drawn uniformly given the observed margins whenever an expected count is below 5
(seeded from the table, so reconstruction stays deterministic).
(seeded from the table, so reconstruction stays deterministic). The G-test
(``test="g"``) applies Yates' continuity correction when the table has one degree
of freedom (two observed symbols), matching
``scipy.stats.chi2_contingency(table, lambda_="log-likelihood")``; transCSSR and
stack CSSR share the same implementation.
``correction="bonferroni"`` divides ``alpha`` by the number of suffixes eligible
for testing, bounding the chance of any spurious split. Because CSSR chooses each
test in light of earlier outcomes, false-discovery-rate step-up procedures do not
Expand Down
4 changes: 3 additions & 1 deletion docs/generators/epsilon_transducer_inference.rst
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,9 @@ channel :cite:`Barnett2015`.
Causal states are equivalence classes of joint ``(input, output)`` pasts that
induce the same conditional next-output law ``P(y | history, x)`` for every input
symbol ``x``. Rare histories inherit their parent's state (controlled by
``min_count``); the split decision uses a G-test at significance ``alpha``.
``min_count``); the split decision uses a G-test at significance ``alpha``, the
same test as :func:`~sofic.generators.epsilon_inference.morphs_differ` (with
Yates' continuity correction at one degree of freedom).
As in :func:`~sofic.generators.epsilon_inference.cssr`, ``test="exact"`` uses a
Monte Carlo exact G-test for tables with small expected counts, and
``correction="bonferroni"`` divides ``alpha`` by the number of
Expand Down
24 changes: 24 additions & 0 deletions docs/generators/hmm_inference.rst
Original file line number Diff line number Diff line change
Expand Up @@ -99,3 +99,27 @@ Score and observed information
.. autofunction:: observed_information
.. autofunction:: free_parameter_labels
.. autofunction:: standard_errors

Transition matrices and start policies
--------------------------------------

.. py:module:: sofic.generators.matrices
:no-index:

Every routine above works with the symbol-labeled joint transition matrices
:math:`T^{(x)}_{ij} = P(S_{t+1} = j, X_t = x \mid S_t = i)` of the model's
Mealy presentation :cite:`Rabiner1989,Ellison2009`, built in one place by
:mod:`sofic.generators.matrices`. The initial state law follows one of two
policies:

- ``"model"`` (likelihoods, decoding, sampling, word probabilities): the
model's ``initial_distribution``, or its stationary distribution when none is
given.
- ``"stationary"`` (block and window statistics such as
``joint_block_distribution`` and directional flow): the stationary process
law, preferring the limit reached from ``initial_distribution`` on reducible
chains.

.. autofunction:: sofic.generators.matrices.start_vector
.. autofunction:: sofic.generators.matrices.emission_tensors
.. autofunction:: sofic.generators.matrices.symbol_matrices
23 changes: 23 additions & 0 deletions sofic/automata/_words.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
"""Shared word enumeration for automata modules."""

from __future__ import annotations

from collections.abc import Iterable, Iterator
from typing import Any

Word = tuple[Any, ...]


def _words_up_to(max_length: int, alphabet: Iterable[Any]) -> Iterator[Word]:
"""Yield every word of length at most ``max_length`` in breadth-first order."""
symbols = tuple(alphabet)
frontier: list[Word] = [()]
yield ()
for _ in range(max_length):
nxt: list[Word] = []
for word in frontier:
for symbol in symbols:
extended = (*word, symbol)
yield extended
nxt.append(extended)
frontier = nxt
97 changes: 53 additions & 44 deletions sofic/automata/active.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@

import numpy as np

from sofic.automata._words import _words_up_to
from sofic.automata.dfa import DFA
from sofic.automata.transducers import MealyMachine

Expand Down Expand Up @@ -135,32 +136,50 @@ def output(self, word: Sequence[Any]) -> Word:
return next(iter(outputs))


def _words_up_to(max_length: int, alphabet: Sequence[Any]) -> Iterator[Word]:
frontier: list[Word] = [()]
yield ()
for _ in range(max_length):
nxt: list[Word] = []
for word in frontier:
for symbol in alphabet:
extended = (*word, symbol)
yield extended
nxt.append(extended)
frontier = nxt
class _BoundedEquivalenceOracle:
"""Approximate equivalence test over a finite stream of candidate words.

Subclasses supply the candidate words (:meth:`_words`) and the comparison of
target versus hypothesis on one word (:meth:`_disagrees`); the first
disagreeing word is the counterexample.
"""

def __init__(self, alphabet: Iterable[Any]) -> None:
self._alphabet = tuple(sorted(alphabet, key=repr))

def _words(self) -> Iterator[Word]:
raise NotImplementedError

def _disagrees(self, hypothesis: Any, word: Word) -> bool:
raise NotImplementedError

def find_counterexample(self, hypothesis: Any) -> Word | None:
for word in self._words():
if self._disagrees(hypothesis, word):
return word
return None


class _MembershipBoundedOracle(_BoundedEquivalenceOracle):
"""Bounded oracle comparing a membership oracle against a DFA hypothesis."""

def __init__(self, membership: MembershipOracle, alphabet: Iterable[Any]) -> None:
super().__init__(alphabet)
self._membership = membership

def _disagrees(self, hypothesis: DFA, word: Word) -> bool:
return self._membership.member(word) != hypothesis.recognizes(word)


class ExhaustiveEquivalenceOracle:
class ExhaustiveEquivalenceOracle(_MembershipBoundedOracle):
"""Bounded exhaustive equivalence test for DFA hypotheses."""

def __init__(self, membership: MembershipOracle, alphabet: Iterable[Any], *, max_length: int = 10) -> None:
self._membership = membership
self._alphabet = tuple(sorted(alphabet, key=repr))
super().__init__(membership, alphabet)
self._max_length = int(max_length)

def find_counterexample(self, hypothesis: DFA) -> Word | None:
for word in _words_up_to(self._max_length, self._alphabet):
if self._membership.member(word) != hypothesis.recognizes(word):
return word
return None
def _words(self) -> Iterator[Word]:
return _words_up_to(self._max_length, self._alphabet)


class AutomatonEquivalenceOracle:
Expand Down Expand Up @@ -206,7 +225,7 @@ def _step(aut: Any, states: frozenset[Any], symbol: Any) -> frozenset[Any]:
return _closure(aut, targets)


class RandomWalkEquivalenceOracle:
class RandomWalkEquivalenceOracle(_MembershipBoundedOracle):
"""Randomized equivalence test drawing random input words for DFA hypotheses."""

def __init__(
Expand All @@ -218,39 +237,33 @@ def __init__(
max_steps: int = 30,
rng: np.random.Generator | int | None = None,
) -> None:
self._membership = membership
self._alphabet = tuple(sorted(alphabet, key=repr))
super().__init__(membership, alphabet)
self._num_walks = int(num_walks)
self._max_steps = int(max_steps)
self._rng = rng if isinstance(rng, np.random.Generator) else np.random.default_rng(rng)

def find_counterexample(self, hypothesis: DFA) -> Word | None:
def _words(self) -> Iterator[Word]:
n_symbols = len(self._alphabet)
for _ in range(self._num_walks):
length = int(self._rng.integers(0, self._max_steps + 1))
word = tuple(self._alphabet[int(self._rng.integers(0, n_symbols))] for _ in range(length))
if self._membership.member(word) != hypothesis.recognizes(word):
return word
return None
yield tuple(self._alphabet[int(self._rng.integers(0, n_symbols))] for _ in range(length))


class MealyExhaustiveEquivalenceOracle:
class MealyExhaustiveEquivalenceOracle(_BoundedEquivalenceOracle):
"""Bounded exhaustive equivalence test for Mealy hypotheses."""

def __init__(self, oracle: MealyMembershipOracle, alphabet: Iterable[Any], *, max_length: int = 10) -> None:
super().__init__(alphabet)
self._oracle = oracle
self._alphabet = tuple(sorted(alphabet, key=repr))
self._max_length = int(max_length)

def find_counterexample(self, hypothesis: MealyMachine) -> Word | None:
for word in _words_up_to(self._max_length, self._alphabet):
if not word:
continue
produced = hypothesis.transduce(word)
hyp_out = next(iter(produced)) if produced else None
if self._oracle.output(word) != hyp_out:
return word
return None
def _words(self) -> Iterator[Word]:
return (word for word in _words_up_to(self._max_length, self._alphabet) if word)

def _disagrees(self, hypothesis: MealyMachine, word: Word) -> bool:
produced = hypothesis.transduce(word)
hyp_out = next(iter(produced)) if produced else None
return self._oracle.output(word) != hyp_out


# ------------------------------------------------------------------------------ L*
Expand Down Expand Up @@ -452,10 +465,6 @@ def build() -> DFA:
raise RuntimeError("TTT did not converge within max_rounds; check the equivalence oracle")


def _hypothesis_access(word: Word, sift: Callable[[Word], _DTNode]) -> Word:
return sift(word).access


def _split_leaf(
counterexample: Word,
sift: Callable[[Word], _DTNode],
Expand All @@ -466,7 +475,7 @@ def _split_leaf(
length = len(counterexample)

def alpha(index: int) -> Word:
return _hypothesis_access(counterexample[:index], sift) + counterexample[index:]
return sift(counterexample[:index]).access + counterexample[index:]

base = member(alpha(0))
breakpoint_index = None
Expand All @@ -477,7 +486,7 @@ def alpha(index: int) -> Word:
if breakpoint_index is None: # pragma: no cover - guaranteed by a valid counterexample
raise RuntimeError("counterexample analysis found no breakpoint")

state_access = _hypothesis_access(counterexample[:breakpoint_index], sift)
state_access = sift(counterexample[:breakpoint_index]).access
symbol = counterexample[breakpoint_index]
discriminator = counterexample[breakpoint_index + 1 :]
new_access = state_access + (symbol,)
Expand Down
40 changes: 20 additions & 20 deletions sofic/automata/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -79,10 +79,6 @@ def intersection(self, other: LabeledAutomaton) -> DFA:

return intersection_dfa(self, other)

def intersect(self, other: LabeledAutomaton) -> DFA:
"""Alias for :meth:`intersection`."""
return self.intersection(other)

def complement(self, alphabet: frozenset[Any] | None = None) -> DFA:
"""Return a complete DFA recognizing the complement over ``alphabet``."""
from sofic.automata.languages.automaton_ops import complement_dfa
Expand All @@ -101,20 +97,12 @@ def concat(self, other: LabeledAutomaton) -> NFA:

return concat_nfa(self, other)

def concatenate(self, other: LabeledAutomaton) -> NFA:
"""Alias for :meth:`concat`."""
return self.concat(other)

def kleene_star(self) -> NFA:
"""Return an NFA recognizing the Kleene star of this language."""
from sofic.automata.languages.automaton_ops import kleene_star_nfa

return kleene_star_nfa(self)

def star(self) -> NFA:
"""Alias for :meth:`kleene_star`."""
return self.kleene_star()

def is_deterministic(self) -> bool:
"""Return whether this automaton is DFA-deterministic."""
from sofic.properties import is_deterministic_automaton
Expand All @@ -140,14 +128,7 @@ def epsilon_closure(self, states: set[Hashable]) -> set[Hashable]:
return closure

def _run_nfa(self, word: Sequence[Any], start: set[Hashable] | None = None) -> set[Hashable]:
seed = set(self.initial_states) if start is None else set(start)
current = self.epsilon_closure(seed)
for symbol in word:
next_states: set[Hashable] = set()
for state in current:
next_states.update(self.delta(state, symbol))
current = self.epsilon_closure(next_states)
return current
return run_nfa(self, word, start=start)

def reverse(self) -> NFA:
"""Return an NFA recognizing the reversed language."""
Expand All @@ -159,3 +140,22 @@ def reverse(self) -> NFA:
accepting_states=frozenset(self.epsilon_closure(set(self.initial_states))),
graph=self.graph.reverse(),
)


def run_nfa(
aut: LabeledAutomaton,
word: Sequence[Any],
start: set[Hashable] | frozenset[Hashable] | None = None,
) -> set[Hashable]:
"""Return the epsilon-closed state set reached by reading ``word``.

Simulation starts from ``start`` (default: ``aut.initial_states``).
"""
seed = set(aut.initial_states) if start is None else set(start)
current = aut.epsilon_closure(seed)
for symbol in word:
next_states: set[Hashable] = set()
for state in current:
next_states.update(aut.delta(state, symbol))
current = aut.epsilon_closure(next_states)
return current
7 changes: 1 addition & 6 deletions sofic/automata/enumeration.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
from itertools import product
from typing import Any

from sofic.automata.algorithms import _effective_alphabet
from sofic.automata.base import LabeledAutomaton


Expand Down Expand Up @@ -36,9 +37,3 @@ def iter_language(
while max_length is None or length <= max_length:
yield from words_of_length(automaton, length)
length += 1


def _effective_alphabet(automaton: LabeledAutomaton) -> frozenset[Any]:
from sofic.automata.algorithms import _effective_alphabet as _shared

return _shared(automaton)
9 changes: 0 additions & 9 deletions sofic/automata/languages/_quotient_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,15 +9,6 @@
from sofic.automata.languages.base import AutomatonLanguage, ExplicitLanguage


def _words_up_to(length: int, alphabet: frozenset[Any]) -> list[tuple[Any, ...]]:
if length < 0:
return []
if length == 0:
return [()]
shorter = _words_up_to(length - 1, alphabet)
return [word + (symbol,) for word in shorter for symbol in alphabet]


def _residual_from_state(aut: AutomatonLanguage, state) -> AutomatonLanguage:
sub = aut.automaton.copy()
sub.initial_states = frozenset({state})
Expand Down
Loading
Loading