Stop engineering prompts. Start declaring contracts.
Flock orchestrates AI agents on a blackboard. Each agent declares which typed artifacts it consumes and which it publishes, and workflows emerge from those subscriptions: there is no graph to wire and no 500-line prompt to maintain. Pydantic models are the contract, independent agents run in parallel on their own, and visibility controls, a real-time dashboard and OpenTelemetry tracing are built in.
New:
- ⚖️ Decision models: route on calibrated decisions instead of generated text.
- ☁️ Microsoft Foundry: run a Flock application as a Foundry hosted agent.
📖 Documentation · Installation · Guides · Examples · Changelog
- Quick start
- What's new: decision models, Microsoft Foundry
- Why Flock
- Features
- Observability
- Examples
- How Flock compares
- Status and roadmap
- Contributing
pip install flock-core # or: uv add flock-core
export OPENAI_API_KEY="sk-..."
export DEFAULT_MODEL="openai/gpt-4.1"import asyncio
from pydantic import BaseModel, Field
from flock import Flock, flock_type
@flock_type
class CodeSubmission(BaseModel):
code: str
language: str
@flock_type
class BugAnalysis(BaseModel):
bugs_found: list[str]
severity: str = Field(pattern="^(Critical|High|Medium|Low|None)$")
@flock_type
class SecurityAnalysis(BaseModel):
vulnerabilities: list[str]
risk_level: str = Field(pattern="^(Critical|High|Medium|Low|None)$")
@flock_type
class FinalReview(BaseModel):
verdict: str = Field(pattern="^(Approve|Approve with Changes|Reject)$")
action_items: list[str]
flock = Flock("openai/gpt-4.1")
# Agents subscribe to types; nobody wires them together
flock.agent("bug_detector").consumes(CodeSubmission).publishes(BugAnalysis)
flock.agent("security_auditor").consumes(CodeSubmission).publishes(SecurityAnalysis)
# AND gate: runs once both analyses are on the blackboard
flock.agent("reviewer").consumes(BugAnalysis, SecurityAnalysis).publishes(FinalReview)
async def main():
await flock.publish(CodeSubmission(code="def add(a, b): return a - b", language="python"))
await flock.run_until_idle()
for review in await flock.store.get_by_type(FinalReview):
print(review.verdict, review.action_items)
asyncio.run(main())- The bug detector and the security auditor run in parallel; the reviewer waits for both.
- There is no prompt: the field names, types and constraints are the instructions.
- Pydantic validates every output, so downstream agents only ever see valid data.
Replace the body of main() with await flock.serve(dashboard=True) to watch the agents in the dashboard and publish artifacts from the browser.
Many agent steps are decisions, not writing tasks: which team handles a ticket, is it urgent, which of 100 controls does a sentence implement. Flock asks decision models for them (Microsoft-Decision-1, OpenAI Decisions, TypeSafe Jev or a local Clef model). They answer a typed question with a calibrated probability for every option in one forward pass, without generating text. The decision is an artifact on the blackboard, and agents subscribe to its answers.
A decision model looks at each shape's image and sorts it into a color lane, live in the dashboard. In-between colors stay below the threshold and land in the UNSURE lane for an inspector (03_color_sorter.py).
from flock import Choice, Flock, YesNo
flock = Flock("openai/gpt-4.1", decision_model="azure/decision-1")
class Route(Choice):
"""Which team should handle this support ticket?"""
billing = "Charges, invoices, refunds"
tech = "Bugs, crashes, login problems"
class Urgent(YesNo):
"""Does the customer need an answer today?"""
flock.agent("triage").consumes(Ticket).decides(Route, Urgent, threshold=0.8) # one request
# Subscribers receive the ticket itself; the decision is on ctx.decisions
flock.agent("billing").consumes(Route.billing).publishes(Reply)
flock.agent("supervisor").consumes(Route.UNSURE).publishes(Reply) # below the threshold
flock.agent("pager").consumes(Route.tech, Urgent.yes).publishes(Page) # AND about one ticket| Question types | Choice, YesNo, Scale (ordered levels, .or_higher) and Checklist (one yes/no per item, e.g. 100 controls, one decision) |
| Large catalogs | Tournaments ask a choice with 1,000+ options in rounds; screening networks split a catalog across parallel nodes and rank the passes |
| Uncertainty | Every decision carries its probabilities; below the threshold it goes to UNSURE, its own branch |
| Providers | azure/, openai/, jev/ and any local /v1/systemone server (local/); images via flock.Image; retries and a shared request budget for rate limits |
| Dashboard | Deciders with their options and counts, edges labelled by option, probability bars, checklist grids and tournament brackets |
Measured with Microsoft-Decision-1 on the repository's examples:
- Routing: 100 arXiv abstracts sorted into research fields at 169 ms per paper, against 598 ms for
gpt-4.1; the decision model agreed with arXiv's category on 90 papers,gpt-4.1on 85. - Compliance: 24 documents checked against 100 controls with one checklist decision each, about 600 ms per document, recall 1.00 at precision 0.75 against ground truth.
📖 Decision models guide · Examples
FlockApplication runs one request as one isolated Flock workflow. It has a typed input, an explicit allowlist of public outputs, results streamed per artifact and exactly one terminal outcome (succeeded, failed, cancelled or timed_out). A queue worker, your own ASGI service or Microsoft Foundry can host it.
from flock import Flock, FlockApplication
from flock.integrations.foundry import FoundryResponsesAdapter
def build_flock() -> Flock: # a fresh blackboard per workflow
flock = Flock("azure/gpt-4.1", no_output=True)
flock.agent("triage").consumes(IncidentRequest).publishes(IncidentTriage)
flock.agent("summarizer").consumes(IncidentTriage).publishes(IncidentSummary)
return flock
application = FlockApplication(
factory=build_flock,
input_type=IncidentRequest,
output_types=(IncidentSummary,),
required_output_types=(IncidentSummary,),
)
host = FoundryResponsesAdapter(
application,
input_mapper=lambda turn: IncidentRequest(report=turn.text),
output_mapper=lambda summary: summary.summary,
)
host.run() # Foundry's Responses protocol on $PORT or 8088- Official host: the adapter uses Microsoft's
azure-ai-agentserver-responseshost. Streaming, background mode, polling and cancellation map onto Flock workflows. - Identity: each Foundry user becomes the workflow's principal. By default, public outputs reach every caller; outputs with
TenantVisibility(tenant_id=...)reach only the user it names, and agent-private outputs reach no caller. - Lean install:
pip install "flock-core[foundry]". Without the extra, no Azure package is installed or imported. - Deployable sample:
examples/14-foundry/incident-triagefollows the official Foundry layout and has been deployed and tested as a hosted agent.
📖 Applications guide · Foundry guide
# Instead of a prompt that explains the output format in 500 lines ...
@flock_type
class BugDiagnosis(BaseModel):
severity: str = Field(pattern="^(Critical|High|Medium|Low)$")
category: str = Field(description="Bug category")
root_cause_hypothesis: str = Field(min_length=50)
confidence_score: float = Field(ge=0.0, le=1.0)
# ... the schema is the instruction
flock.agent("diagnostician").consumes(BugReport).publishes(BugDiagnosis)Schemas survive model upgrades, fail loudly at parse time instead of in production, and can be tested with concrete inputs and outputs.
# Graph frameworks: every new agent means new edges
workflow.add_edge("radiologist", "diagnostician")
workflow.add_edge("lab_tech", "diagnostician")
# Flock: agents subscribe to types, and the workflow emerges
flock.agent("radiologist").consumes(Scan).publishes(XRayAnalysis)
flock.agent("lab_tech").consumes(Scan).publishes(LabResults)
flock.agent("diagnostician").consumes(XRayAnalysis, LabResults).publishes(Diagnosis)
flock.agent("perf").consumes(Scan).publishes(PerfAnalysis) # added later, nothing rewiredAgents know data types, not each other. Independent agents run concurrently, an agent with several inputs waits for all of them, and adding an agent is one subscription. Blackboard systems have coordinated specialists since the 1970s (Hearsay-II, HASP/SIAP, BB1); Flock applies the pattern to LLM agents.
flock.agent("diagnostician").consumes(XRayAnalysis, LabResults) # AND: wait for both
flock.agent("alerts").consumes(SystemAlert).consumes(UserAlert) # OR: either one
flock.agent("aggregator").consumes(Order, Order, Order) # three Orders
flock.agent("urgent").consumes(Diagnosis, where=lambda d: d.severity == "Critical")
flock.agent("payments").consumes(Transaction, batch=BatchSpec(size=25, timeout=timedelta(seconds=30)))
flock.agent("notify").consumes(
Order, Shipment, join=JoinSpec(by=lambda x: x.order_id, within=timedelta(hours=24))
)📖 Predicates · Batching · Joins
One execution can publish many artifacts, including several types at once:
flock.agent("ideas").consumes(Brief).publishes(Idea, fan_out=(5, 20), where=lambda i: i.score >= 8)
flock.agent("studio").consumes(Idea).publishes(Movie, MovieScript, MovieCampaign, fan_out=3) # 9 artifacts, one callRoute by meaning with local embeddings (all-MiniLM-L6-v2 on ONNX via fastembed, no external API):
# pip install "flock-core[semantic]"
flock.agent("security").consumes(Ticket, semantic_match="security vulnerability exploit").publishes(Alert)
flock.agent("billing").consumes(Ticket, semantic_match="payment charge refund", semantic_threshold=0.6)Semantic subscriptions compare embeddings locally. Decision models give calibrated probabilities over a closed set of options. The decision models guide explains when to use which.
📖 Semantic subscriptions guide
flock.agent("health").schedule(every=timedelta(seconds=30)).publishes(HealthStatus)
flock.agent("daily").schedule(at=time(hour=17)).publishes(DailyReport) # 17:00 UTC
flock.agent("workdays").schedule(cron="0 9 * * 1-5").publishes(WorkdayReport) # cron in UTC
flock.agent("errors").schedule(every=timedelta(minutes=5)).with_context(ErrorLogs()).publishes(ErrorReport)A timer run sees the blackboard as context; a context provider such as ErrorLogs() above narrows it. Engines can read ctx.trigger_type == "timer", ctx.timer_iteration and ctx.fire_time.
Access control is part of the blackboard, not an add-on. It decides which agents an artifact triggers and which artifacts they see:
agent.publishes(CustomerData, visibility=TenantVisibility(tenant_id="customer_123"))
agent.publishes(MedicalRecord, visibility=PrivateVisibility(agents={"physician", "nurse"}))
agent.publishes(IntelReport, visibility=LabelledVisibility(required_labels={"clearance:secret"}))
flock.agent("analyst").labels("clearance:secret").consumes(IntelReport)Context providers shape what an agent sees beyond its trigger, for example only recent, correlated or tagged artifacts. The built-in providers extend BaseContextProvider, which applies visibility filtering. Custom providers that extend it get the same guarantee; a provider that implements the protocol directly must filter by visibility itself.
from flock import FilterConfig
from flock.core.context_provider import FilteredContextProvider
flock = Flock("openai/gpt-4.1", context_provider=FilteredContextProvider(FilterConfig(tags={"urgent"})))📖 Visibility · Context providers
The blackboard is in memory by default. Two persistent stores come with Flock:
- SQLite keeps a full, queryable history:
Flock(..., store=SQLiteBlackboardStore(".flock/blackboard.db"))fromflock.core.store. - Dapr state stores (Redis, PostgreSQL, Cosmos DB and others) let several Flock processes share one blackboard. Install it with
pip install "flock-core[dapr]".
📖 Persistent blackboard · Dapr state store
- Agent components hook into every step of an agent run:
on_pre_consume,on_pre_evaluate,on_post_evaluateand more. - Orchestrator components see the whole blackboard; scheduling and webhooks are built this way.
- Server components compose the HTTP API.
- Engines decide how an agent turns inputs into outputs: DSPy (the default), decision models or your own logic.
- Built-in safeguards:
- a circuit breaker stops runaway cascades (1,000 iterations by default);
- agents do not trigger themselves (
prevent_self_trigger); - duplicate deliveries are filtered;
.best_of(n, score=...)picks the best of several runs.
- Azure OpenAI with Entra ID:
DSPyEngine(lm_kwargs={"azure_ad_token_provider": get_default_azure_token_provider()}), with theazureextra. - MCP servers are available as agent tools.
📖 Agent components · Orchestrator components · DSPy engine and Azure · Local models
await flock.serve(dashboard=True) starts the REST API and the dashboard on port 8344.
- Dashboard:
- an agent view and a blackboard view with live updates over WebSocket;
- five auto-layouts;
- filters by correlation id and time;
- publish artifacts and run agents from the browser.
- Trace viewer: timeline, statistics, RED metrics, dependencies and configuration.
- Tracing:
FLOCK_AUTO_TRACE=trueandFLOCK_TRACE_FILE=truerecord every operation with its inputs and outputs as OpenTelemetry spans in.flock/traces.duckdb. Query them with SQL, or export them via OTLP. - REST API:
POST/GET /api/v1/artifacts,POST /api/v1/agents/{name}/run,GET /api/v1/correlations/{id}/status,/healthand/metrics;- OpenAPI docs at
/docs.
📖 Dashboard · Tracing · REST API
git clone https://github.com/whiteducksoftware/flock.git && cd flock
uv sync --all-extras
uv run python examples/01-getting-started/01_declarative_pizza.py| Folder | What it shows |
|---|---|
01-getting-started |
First agents, inputs and outputs, MCP, tracing, joins, batches, webhooks |
02-patterns |
Publishing (fan-out, multi-output), visibility, complex patterns |
03-hackathon |
A progressive hands-on tutorial |
04-misc |
Persistent blackboard, dashboard edge cases, a 100-agent scale test |
05-engines – 07-orchestrator-components |
Custom engines, agent components and orchestrator components |
08-semantic |
Semantic ticket routing and filtering |
09-server-components |
Every server component, and a full composition |
10-scheduling |
Interval, daily, one-time and cron timers |
12-dapr |
Dapr state stores (in-memory, PostgreSQL, encrypted Redis) |
13-applications |
FlockApplication behind a worker and a Starlette SSE endpoint |
14-foundry |
A deployable Microsoft Foundry hosted agent |
15-decisions |
Decision models: triage, arXiv race, image sorting, question types, checklists, tournaments |
| Graph-based | Chat-based | Flock (blackboard) | |
|---|---|---|---|
| Coordination | Hand-wired edges | Message passing | Type subscriptions |
| Parallelism | Manual split/join | Mostly sequential | Automatic |
| Outputs | Varies | Text messages | Validated Pydantic models |
| Adding an agent | Rewire the graph | Update the flow | One subscription |
| Routing | Code or an LLM's text | The conversation | Types, predicates, semantic matches and calibrated decisions |
| Access control | DIY | DIY | Built-in visibility |
| Testing | The whole graph | The whole group | Each agent in isolation |
Flock fits when you want typed outputs, parallel agents without wiring, routing you can test, and access control and tracing from the start.
Look elsewhere when your workflow is a strictly sequential script, when you need a large ecosystem of ready-made integrations, or when your team is invested in another framework. Flock's community is smaller than those of the big graph and chat frameworks.
Flock 0.5 is the current release line (flock-core on PyPI), with about 2,800 tests. The core is stable:
- the blackboard with typed contracts, subscriptions and visibility;
- the dashboard and tracing;
- the SQLite and Dapr stores;
- decision models;
- application hosting, including Microsoft Foundry.
The work toward 1.0 is tracked in issues labelled [1.0]:
- bounded admission and concurrency, and execution budgets;
- run history that survives restarts, and run cancellation;
- authenticated dashboard and WebSocket access;
- explanations for why a subscription did not trigger;
- agent skills.
ROADMAP.md summarizes the planned work: the Flock 1.0 scope and separately qualified integrations, both for Q4 2026.
Flock is developed in the open. Please read the contributing guide first.
- AI coding agents (and humans who like checklists) should start with AGENTS.md.
- The architecture overview, error handling and async patterns describe the conventions the code follows.
- Tests must pass and code is formatted with Ruff (
uv run poe test,uv run poe lint).
MIT, see LICENSE. Built with ❤️ by white duck GmbH.




