Control¶
The admission gate: decide whether a consequential action may run, and record the decision either way.
tulip.control is the domain-neutral surface, and the import path to use.
Older releases shipped these under tulip.security; that path still resolves,
with a deprecation warning, until 3.0.
For the concepts, start with The control layer and Writing a policy that holds.
Admitting an action¶
admit() evaluates the policy, records the decision on the audit trail, and
runs the action only if it was allowed. A held or denied action raises
AdmissionError carrying the ApprovalDecision that explains why.
admit
async
¶
admit(action: Action, perform: Callable[[], Awaitable[T]], *, policy: ControlPolicy, finding: Evidence | None = None, verdict: VerificationResult | None = None, trail: AuditTrail | None = None, approved_by: str | None = None, ledger: SpendLedger | None = None, spend_scope: str = 'default') -> T
Run perform only if action clears the trust chain; else reject.
The mandatory gate that turns the composable chain into an enforced one:
- :func:
~tulip.control.policy.approveweighs the action against the evidence (finding), the verification (verdict), and thepolicy. - The decision is recorded to
trail(if given) — admitted or not — so no side effect is un-audited. - On ALLOW,
performis awaited and its result returned. On require_human or deny, :class:AdmissionErroris raised with the decision attached.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
action
|
Action
|
The proposed side-effecting action. |
required |
perform
|
Callable[[], Awaitable[T]]
|
A zero-arg async callable that performs the side effect. |
required |
policy
|
ControlPolicy
|
The governing :class: |
required |
finding
|
Evidence | None
|
The evidence the action responds to. |
None
|
verdict
|
VerificationResult | None
|
The :func: |
None
|
trail
|
AuditTrail | None
|
An :class: |
None
|
approved_by
|
str | None
|
Who approved a |
None
|
ledger
|
SpendLedger | None
|
Where |
None
|
spend_scope
|
str
|
The scope the spend counts against: a customer, a tenant. |
'default'
|
Returns:
| Type | Description |
|---|---|
T
|
Whatever |
Raises:
| Type | Description |
|---|---|
AdmissionError
|
if the action is not admitted (deny, or require_human
without |
Source code in .sdk/src/tulip/control/admission.py
AdmissionError ¶
Bases: Exception
A side-effecting action failed admission — it did not clear the trust chain.
Carries the :class:~tulip.control.policy.ApprovalDecision so the caller can
route a require_human hold to an approver or surface a deny reason.
Source code in .sdk/src/tulip/control/admission.py
Deciding¶
approve() is the pure decision function — no I/O, no side effects. It takes
an action and a policy and returns the outcome. Rules combine by taking the
strongest result, so deny beats hold beats allow (in the API,
ApprovalOutcome.DENY > REQUIRE_HUMAN > ALLOW).
approve ¶
approve(action: Action, *, policy: ControlPolicy, finding: Evidence | None = None, verdict: VerificationResult | None = None, advisor: ControlAdvisor | None = None, spent_usd: float = 0.0) -> ApprovalDecision
Decide whether action may proceed: allow / require_human / deny.
Weighs every rule and returns the strongest triggered outcome (deny > require_human > allow), recording each check that fired so the decision is auditable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
action
|
Action
|
The proposed action. |
required |
policy
|
ControlPolicy
|
The governing :class: |
required |
finding
|
Evidence | None
|
The evidence the action responds to (optional). |
None
|
verdict
|
VerificationResult | None
|
The :func: |
None
|
spent_usd
|
float
|
What the action's scope has already spent, for
|
0.0
|
advisor
|
ControlAdvisor | None
|
An optional trained control model. It may only raise the
decision toward caution — see :func: |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
An |
ApprovalDecision
|
class: |
Source code in .sdk/src/tulip/control/policy.py
142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 | |
ControlPolicy
dataclass
¶
ControlPolicy(require_verification_score: float = 0.8, max_blast_radius: int = 1, require_human_for: frozenset[str] = (lambda: frozenset({'production'}))(), deny_for: frozenset[str] = frozenset(), min_severity: Severity = Severity.LOW, require_sandbox_for: frozenset[str] = frozenset(), version: str = '', require_human_over_usd: float | None = None, spend_limit_usd: float | None = None)
The CISO knobs. Defaults are conservative — auto-allow only the safe path.
require_verification_score: minimum :class:VerificationResultconfidence to auto-allow; below it (or with no verdict) a human is required.max_blast_radius: most assets an action may affect to auto-allow.require_human_for: action labels (environment / kind / tag) that always need a human (default: anything inproduction).deny_for: labels that are hard-denied outright.min_severity: don't act on findings below this band.require_sandbox_for: labels whose actions must execute in a sandbox — an action matching one of these is denied unless it carries the :data:SANDBOXED_TAGtag. Enforced at the agent loop's tool seam by :class:~tulip.tools.sandbox.SandboxEnforcerHook.version: a label for this policy. When set it is bound into approval ids, so a decision made under one version is never redeemed under another.require_human_over_usd: an action whosecost_usdis above this needs a human.spend_limit_usd: an action that would take its scope's cumulative spend past this is denied — no approval overrides it. The caller supplies what the scope has spent (spent_usd), typically from a spend ledger.
ApprovalDecision
dataclass
¶
ApprovalDecision(outcome: str, reason: str, action: Action, checks: list[str] = list(), policy_outcome: str = ApprovalOutcome.ALLOW, model_outcome: str | None = None)
The outcome of weighing an action against evidence, verification, and policy.
escalated_by_model
property
¶
Whether a control model made this decision stricter than policy alone.
ApprovalOutcome ¶
Outcome labels (kept simple/stable as plain strings).
Describing an action¶
A policy matches on what an action is — its environment, kind, blast radius, and tags — never on the name of the tool performing it.
Action
dataclass
¶
Action(name: str, asset: str = '', blast_radius: int = 1, environment: str = 'unknown', kind: str = '', tags: frozenset[str] = frozenset(), cost_usd: float = 0.0)
A proposed response action, with the attributes policy reasons over.
labels ¶
Gating a tool¶
gate_tool puts the gate in front of a tool the agent already has. The
returned tool keeps the original's name, description and parameter schema, so
the model sees no difference and nothing else in the agent changes — which is
what makes the control structural rather than advisory. There is nothing to
notice, so nothing to talk around.
from tulip.control import ControlPolicy, gate_tool
agent = Agent(model=model, tools=[
lookup_order, # read-only, ungated
gate_tool(issue_refund, policy=ControlPolicy()), # gated
])
A refusal comes back to the model as a readable result naming the outcome and
the reason, so the agent can explain the hold rather than the run ending in a
traceback. That is on_refusal="return", the default. Pass
on_refusal="raise" for a caller that would rather stop, or
on_refusal="interrupt" to pause the run on a hold until someone decides; a
denial still comes back as a refusal (see
Pausing until someone decides).
Gating a sandboxed tool composes rather than replacing it: the gate decides, and only an admitted call reaches the sandbox.
gate_tool ¶
gate_tool(tool: Tool, *, policy: ControlPolicy, action: ActionSpec | None = None, trail: AuditTrail | None = None, finding: Evidence | None = None, verdict: VerificationResult | None = None, on_refusal: Literal['return', 'raise', 'interrupt'] = 'return', approval: ApprovalBridge | None = None, principal: str = 'agent', refusal_reason: str | Callable[[ApprovalDecision], str] | None = None, approval_context: Mapping[str, str] | Callable[[str, dict[str, Any]], Mapping[str, str]] | None = None, ledger: SpendLedger | None = None, spend_scope: str | Callable[[str, dict[str, Any]], str] = 'default') -> Tool
Return a copy of tool whose call goes through :func:admit first.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tool
|
Tool
|
The tool to gate. Not modified — an ungated reference stays usable, which matters when the same function is called by trusted code elsewhere. |
required |
policy
|
ControlPolicy
|
The :class: |
required |
action
|
ActionSpec | None
|
How to turn a call into an :class: |
None
|
trail
|
AuditTrail | None
|
Records every decision, allowed or not. Omit and decisions are weighed but not written down. |
None
|
finding
|
Evidence | None
|
Grounded evidence supporting the action, when the policy requires one. |
None
|
verdict
|
VerificationResult | None
|
A verification result, when the policy sets
|
None
|
approval
|
ApprovalBridge | None
|
Where to submit an action held for a human. Without one, a
hold tells the model it was held and stops there — true, and not
actionable. With one, the refusal carries an |
None
|
principal
|
str
|
Who the held action is attributed to on the approval. |
'agent'
|
approval_context
|
Mapping[str, str] | Callable[[str, dict[str, Any]], Mapping[str, str]] | None
|
What an approval is bound to besides the call itself,
as a mapping or |
None
|
ledger
|
SpendLedger | None
|
A spend ledger. The gate reads the scope's cumulative spend
before each decision, for |
None
|
spend_scope
|
str | Callable[[str, dict[str, Any]], str]
|
The scope spend counts against, as a string or
|
'default'
|
on_refusal
|
Literal['return', 'raise', 'interrupt']
|
|
'return'
|
refusal_reason
|
str | Callable[[ApprovalDecision], str] | None
|
What the model is told when an action is refused. By
default it is the policy's own reason, which names the checks that
fired — accurate, and written in control-plane vocabulary the model
will repeat verbatim to the end user ("blast radius 3 exceeds the
maximum 1"). Pass a string, or |
None
|
Returns:
| Type | Description |
|---|---|
Tool
|
A new :class: |
Tool
|
parameter schema are the original's, so it is a drop-in wherever the |
Tool
|
original was passed — the model cannot tell the difference, which is |
Tool
|
the point: the gate is not something the model can be talked around. |
"return" is the default because a refusal is information the agent can
act on. An exception ends the run, and "the refund was held for a human" is
something the user should hear rather than a stack trace.
Source code in .sdk/src/tulip/control/gate.py
161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 | |
Holding an action for a human¶
A hold is only useful if the agent can find out what happened next. Give
gate_tool an approval bridge and a held refusal carries an approval_id the
agent can poll, while a human decides on a channel the agent cannot reach:
{"status": "held_for_approval", "outcome": "require_human",
"action": "issue_refund", "asset": "ord-4821", "reason": "...",
"approval_id": "appr-77",
"next": "call approval_status(approval_id) once a human decides"}
A denial deliberately gets no id. It is final, and offering one would invite the agent to wait for a decision that is not coming.
ApprovalBridge is a structural Protocol with no import-time dependency, so
any approval queue with submit and state methods satisfies it. The two
stores Tulip ships, InMemoryApprovals and FileApprovals, satisfy it too;
see Pausing until someone decides.
ApprovalBridge ¶
Bases: Protocol
Submit a held action for out-of-band approval, and check its state.
A structural Protocol, deliberately: it has no import-time dependency on anything, so any approval queue with these two methods satisfies it.
Without one, a held action tells the model it was held and stops there — true, and not actionable. With one, the refusal carries an id the agent can poll while a human decides on a channel the agent cannot reach.
submit ¶
What the user hears when an action is refused¶
The reason in that payload is, by default, the policy's own — a join of the
checks that fired:
That is the right level of detail for the audit trail and for a developer
reading a log. It is also control-plane vocabulary, and a model handed it
repeats it verbatim. Run against a live model, the refusal above reached the
customer as "the blast radius (3) exceeds the maximum 1" and "it's
classified as a large_refund".
refusal_reason gives the model the sentence you want the user to hear
instead. Pass a string, or (decision) -> str to vary it by outcome:
gate_tool(
issue_refund,
policy=policy,
trail=trail,
refusal_reason=lambda d: (
"We can't refund this amount automatically."
if d.outcome == "deny"
else "This refund is waiting on a manager."
),
)
The full policy reason still goes to the trail — a friendlier sentence for the customer must not shrink the record. Added in 2.10.0.
Pausing until someone decides¶
A polled id keeps the run going while a human decides. With
on_refusal="interrupt" and an approval store, a hold pauses the run instead:
the agent yields an InterruptEvent whose metadata carries the
approval_id, the checkpointer keeps the conversation, and the store keeps the
pending approval. A person decides with decide(), and
agent.resume(..., thread_id=..., perform_dangling=True) re-issues the held
call, which finds the decision:
from tulip.control import FileApprovals, gate_tool
store = FileApprovals("approvals.json")
refund = gate_tool(
issue_refund,
policy=policy,
approval=store,
on_refusal="interrupt",
trail=trail,
)
agent = Agent(model=model, tools=[refund], checkpointer=FileCheckpointer("checkpoints"))
async for event in agent.run("refund order 4821", thread_id="t1"):
if isinstance(event, InterruptEvent):
approval_id = event.metadata["approval_id"] # the run is parked
# Later, from any process that can open the same file:
store.decide(approval_id, "approved", by="[email protected]")
async for event in agent.resume("approved", thread_id="t1", perform_dangling=True):
...
An approval names one call. Its id is derived from a digest of the principal,
the tool, the arguments, policy.version when set, and any approval_context,
so a call with different arguments waits for its own decision. An approved call
is weighed against the policy again when it is redeemed, so a deny such as a
spend limit still refuses it, and it runs at most once as long as the store's
writes do not race (see the FileApprovals limit below): the approval is
consumed immediately before the side effect, so a repeated call holds again
instead of riding an old yes. A denied call returns a refusal. A policy deny never pauses. on_refusal="interrupt"
needs an ApprovalStore, not a bare ApprovalBridge, and raises TypeError
without one.
InMemoryApprovals lives in one process and is gone on restart; it is for
tests and demos. FileApprovals keeps every record in one JSON file, re-read on
every call and written atomically, so a decision written by another process is
seen by the next call. Writes are serialised only through one FileApprovals
object in one process, and the agent's own submits and consumes are writes too.
A decision saved from another process, or through another FileApprovals
object, at the same moment can overwrite one of them. In the worst case a
consumed approval goes back to approved and can be redeemed again. Where that
matters, keep every writer in one process and give them the same FileApprovals
object, or implement ApprovalStore over a database with atomic updates.
Who may decide. Without an ApprovalAuthority, any named principal can
decide. With one, every decision is checked when it is made:
from tulip.control import ApprovalAuthority, ApproverRule, FileApprovals
authority = ApprovalAuthority(
rules=(
ApproverRule(labels=frozenset({"payment"}), roles=frozenset({"finance"})),
ApproverRule(
labels=frozenset({"production"}),
approvers=frozenset({"olga", "sam"}),
quorum=2,
),
),
roles_of=directory.roles_for,
)
store = FileApprovals("approvals.json", authority=authority)
Every rule that matches the held action's labels must reach its quorum of
distinct authorised approvers before the call is approved; one authorised
denial ends it. The principal that requested the action cannot approve it,
directly or through a delegation it granted, unless every matching rule sets
allow_self_approval. A Delegation lends an approver's authority to someone
else until a deadline. A decision by someone without authority raises
ApprovalAuthorityError and is kept on the record as a rejection. An action no
rule covers cannot be approved by anyone: the authority fails closed.
ApprovalStore ¶
Bases: Protocol
Where held calls wait for a decision.
A superset of :class:~tulip.control.ApprovalBridge: submit and
state keep that shape, so a store works anywhere a bridge does.
submit ¶
submit(principal: str, tool: str, args: Mapping[str, Any], *, reason: str = '', labels: Iterable[str] = (), context: Mapping[str, str] | None = None) -> str
Record a held call, or return the live record for the same call.
Source code in .sdk/src/tulip/control/approvals.py
state ¶
get ¶
decide ¶
decide(approval_id: str, verdict: Verdict, *, by: str, arguments: Mapping[str, Any] | None = None) -> ApprovalRecord
Approve or deny a pending record, naming who decided.
consume ¶
InMemoryApprovals ¶
Bases: _Approvals
An approval store in this process only. Gone on restart; for tests and demos.
Source code in .sdk/src/tulip/control/approvals.py
FileApprovals ¶
Bases: _Approvals
An approval store in one JSON file.
Every call re-reads the file, so a decision written by another process is seen by the next call. Writes are atomic (temporary file, then rename), so a crash never leaves a half-written store. Concurrent writers are serialised within a process only: use one process to decide, or a database-backed store for many.
Source code in .sdk/src/tulip/control/approvals.py
ApprovalRecord
dataclass
¶
ApprovalRecord(approval_id: str, digest: str, principal: str, tool: str, arguments: dict[str, Any], status: Status = 'pending', reason: str = '', created_at: str = _now(), verdict: Verdict | None = None, decided_by: str | None = None, decided_at: str | None = None, labels: list[str] = list(), approvals: list[dict[str, Any]] = list(), rejections: list[dict[str, Any]] = list(), context: dict[str, str] = dict(), approved_arguments: dict[str, Any] | None = None)
One held call and what became of it.
approvers
property
¶
Distinct principals whose approval was accepted, in order.
call_digest ¶
call_digest(principal: str, tool: str, arguments: Mapping[str, Any], context: Mapping[str, str] | None = None) -> str
SHA-256 over the canonical form of one call.
Key order does not matter; any change to a value, the tool, the principal
or the context does. context binds the call to where it was made, such
as a policy version or a thread. An empty context hashes exactly like no
context, so records written before contexts existed still match.
Source code in .sdk/src/tulip/control/approvals.py
ApprovalAuthority
dataclass
¶
ApprovalAuthority(rules: tuple[ApproverRule, ...], delegations: tuple[Delegation, ...] = (), roles_of: Callable[[str], Iterable[str]] | None = None, clock: Callable[[], datetime] = _utcnow)
Approver rules, delegations, and how to look up a principal's roles.
Rules are evaluated by position, so keep their order stable for records that are still pending.
matching ¶
The rules governing an action carrying labels, by position.
Source code in .sdk/src/tulip/control/approvals.py
check ¶
Which matching rules by may decide under, and on what basis.
Raises:
| Type | Description |
|---|---|
ApprovalAuthorityError
|
No rule covers the action, |
Source code in .sdk/src/tulip/control/approvals.py
satisfied ¶
Whether every matching rule has its quorum of distinct approvers.
Source code in .sdk/src/tulip/control/approvals.py
ApproverRule
dataclass
¶
ApproverRule(labels: frozenset[str] = frozenset(), approvers: frozenset[str] = frozenset(), roles: frozenset[str] = frozenset(), quorum: int = 1, allow_self_approval: bool = False)
Who may decide actions that carry some labels.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
labels
|
frozenset[str]
|
Labels this rule governs. Empty matches every action. |
frozenset()
|
approvers
|
frozenset[str]
|
Principals allowed to decide, by name. |
frozenset()
|
roles
|
frozenset[str]
|
Roles allowed to decide, resolved per principal by
:attr: |
frozenset()
|
quorum
|
int
|
Distinct authorised approvals needed. |
1
|
allow_self_approval
|
bool
|
Whether the principal that requested the action may approve it. Off by default: separation of duties. |
False
|
matches ¶
Delegation
dataclass
¶
An approver lends their authority to someone else until a deadline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
grantor
|
str
|
The principal whose authority is lent. |
required |
grantee
|
str
|
The principal who may use it. |
required |
expires_at
|
str
|
ISO-8601 deadline; a naive timestamp is read as UTC. |
required |
labels
|
frozenset[str]
|
Limit the delegation to actions carrying these labels. Empty lends everything the grantor may decide. |
frozenset()
|
active ¶
Whether the deadline is still in the future.
Source code in .sdk/src/tulip/control/approvals.py
ApprovalAuthorityError ¶
Bases: PermissionError
A decision was refused because the decider lacks authority for it.
Capping spend¶
Action.cost_usd says what an action spends. ControlPolicy can hold an
action above a per-action cost (require_human_over_usd) and deny one that
would take its scope's cumulative spend past a limit (spend_limit_usd); no
approval overrides that denial. A spend ledger supplies the cumulative figure:
from tulip.control import Action, ControlPolicy, FileSpendLedger, gate_tool
refund = gate_tool(
issue_refund,
policy=ControlPolicy(spend_limit_usd=5_000, require_human_over_usd=500),
action=lambda name, args: Action(
name=name, asset=args["order_id"], cost_usd=args["amount_usd"]
),
ledger=FileSpendLedger("spend.json"),
spend_scope=lambda name, args: f"customer:{args['customer_id']}",
)
A scope is any string: a thread, a customer, a tenant, a month. Spend is
recorded only after the action has run, so a refused or failed action costs
nothing. The check and the record are not one transaction. Two calls against
the same scope that overlap can each pass the check before either records, so
the cap can be exceeded. That includes two processes, and also two tool calls
the agent runs together in one turn (the default is
tool_execution="concurrent"). SpendLedger reads and records in separate
calls, before and after the action, so a ledger cannot make the check atomic by
itself. Where the cap must hold, run the agent that has the spending tools with
tool_execution="sequential" and keep one writer per scope.
InMemorySpendLedger is the in-process version, for tests
and demos. admit() takes ledger= and a string spend_scope= directly.
InMemorySpendLedger ¶
FileSpendLedger ¶
A spend ledger in one JSON file, with every entry kept.
Re-read on every call and written atomically (temporary file, then rename),
like :class:~tulip.control.FileApprovals. Writers are serialised within a
process only.
Source code in .sdk/src/tulip/control/spend.py
entries ¶
Deriving action labels¶
Turn a tool call into an Action using declarative rules, so the labels a
policy matches on are not hand-written per call site.
resolve_action ¶
Resolve an :class:ActionSpec (or None) into a concrete :class:Action.
Source code in .sdk/src/tulip/control/action.py
default_action ¶
default_action(name: str, kwargs: Mapping[str, Any], *, environment: str = 'unknown', kind: str = '', blast_radius: int = 1, tags: frozenset[str] | None = None) -> Action
A conservative :class:Action for name when none was supplied.
Fail-safe by construction: environment="unknown" plus the stock
:class:~tulip.control.policy.ControlPolicy (which requires a verification
score) lands an un-verified call on require_human rather than
auto-allowing it.
tags defaults to the action's own name, so a policy can always gate one
specific tool by naming it — the one thing that worked before labels were
derived at all.
Source code in .sdk/src/tulip/control/action.py
action_from_labels ¶
action_from_labels(name: str, kwargs: Mapping[str, Any], *, labels: Mapping[str, Any] | None = None, environment: str | None = None, blast_radius: int = 1) -> Action
Build an :class:Action from a tool's declared labels.
labels is what a tool definition declares about the actions it performs
— environment, kind, blast_radius, tags. Anything absent
falls back: the caller's environment (the agent's, or the deployment's),
then "unknown".
The tool's own name is always among the tags, so naming a tool in
require_human_for keeps working regardless of what it declares.
labels["derive"] may carry argument-derived rules (see
:func:derive_labels) — the only part of this that reads kwargs for
labelling. Derived tags join the declared ones, a derived set_kind /
set_environment wins over the declared value (it describes this call),
and a derived blast radius only ever raises the declared one. With no
derive key the result is exactly what it was before.
Source code in .sdk/src/tulip/control/action.py
derive_labels ¶
Evaluate a tool's derive rules against one call's arguments.
Declarative and total: comparisons only, never eval, never a callable,
so nothing a tool receives can execute during labelling. Rules apply in
order and all matching rules apply. Anything that cannot be evaluated — a
missing argument, an argument of the wrong type, a malformed rule — is
skipped and records :data:UNDETERMINED_TAG, so "we could not tell"
reaches the policy as a fact rather than as silence.
Source code in .sdk/src/tulip/control/action.py
DerivedLabels ¶
Accumulator for what a derive list adds to an action.
Source code in .sdk/src/tulip/control/action.py
raise_radius ¶
Deriving may raise the blast radius; it may never lower it.
asset_from_args ¶
Best-effort asset label from a tool call's arguments.
The record¶
A hash-chained log of every decision. Each record commits to the previous
hash, so editing any record breaks verify().
Tamper-evident, not tamper-proof
By default this is a keyless SHA-256 chain held in memory. It detects edits when checked against a head hash you retain out-of-band; it does not prevent them, and it does not anchor the log. Without that head, anyone who can write the log can rebuild an unsigned chain around an edit. Persist the JSONL and pin the head hash externally before relying on it as compliance evidence, and sign the trail to catch a rebuild.
What verify() catches, and the one thing it cannot¶
Called with no arguments, verify() catches every edit, every reorder, and
every deletion from the middle of the chain — each leaves a record whose
stored hash no longer matches its contents, or whose prev_hash no longer
points at the record before it.
It cannot, on its own, catch a truncation:
| Attack | verify() |
verify(expected_head=…) |
|---|---|---|
| Edit a record | False |
False |
| Reorder records | False |
False |
| Delete from the middle | False |
False |
| Drop records off the end | True |
False |
| Discard the trail entirely | True |
False |
Dropping the tail leaves a shorter chain that is perfectly valid on its own
terms. That is a property of hash chains in general, not of this
implementation: nothing inside a chain can attest to a link that was never
handed to it. An agent that can reach its own audit trail can therefore erase
the end of it and still pass a bare verify().
Anchoring closes it. Every attack in that table moves the head, so keep
head somewhere the agent cannot reach — a WORM bucket, an append-only log,
a co-signer, a transparency log — and pass it back:
trail = AuditTrail()
...
anchor = trail.head # to durable, external storage
# later, on the exported chain
restored = AuditTrail.from_records(records)
restored.verify(expected_head=anchor) # False if anything was removed
Added in 2.10.0, alongside a correction: verify() previously documented
itself as detecting "no edit, deletion, or reorder", which overstated what a
chain can prove about its own tail.
Signing the trail¶
A hash chain proves the records agree with each other, not who wrote them.
Rewrite a record and recompute every hash after it, and a bare verify()
passes again. Give the trail a signer and every record's hash is signed with
Ed25519, and the record carries the key_id of the key that signed it:
from tulip.control import AuditTrail, Ed25519Signer, verify_jsonl
trail = AuditTrail(signer=Ed25519Signer.from_pem(private_pem, key_id="audit-2026-09"))
...
anchor = trail.head # to durable, external storage
exported = trail.export_jsonl()
# An auditor with the export and the public key, and no Tulip runtime state:
verify_jsonl(exported, keys={"audit-2026-09": public_pem}, expected_head=anchor)
With keys, an unsigned record, an unknown key id, or a bad signature fails
verification, so a chain rebuilt around an edit fails unless it was signed with
a key the verifier trusts. trail.verify(keys=...) runs the same check in
process.
Signing does not catch truncation. Records dropped off the end take their
signatures with them, and what is left is still validly signed, so keep passing
an externally held expected_head.
To rotate keys, call trail.use_signer(new_signer). Records already written
keep the key they were signed with, so give the verifier both public keys.
Signing needs the cryptography package (pip install "tulip-agents[audit]");
an unsigned trail needs nothing and exports exactly as before.
AuditTrail ¶
An append-only, hash-chained log of agent actions.
Append with :meth:record (or :meth:record_event for a Tulip event);
check integrity with :meth:verify; ship with :meth:export_jsonl.
Pass clock to make timestamps deterministic in tests, and signer
to sign every record.
Source code in .sdk/src/tulip/control/audit.py
use_signer ¶
Sign records from now on with signer — how a key is rotated.
Records already written keep the key they were signed with; a verifier given both public keys accepts the whole trail.
Source code in .sdk/src/tulip/control/audit.py
record ¶
Append a record committing to the current chain head.
Source code in .sdk/src/tulip/control/audit.py
record_event ¶
Append a record for a Tulip event (duck-typed; safe scalar fields).
Source code in .sdk/src/tulip/control/audit.py
records ¶
verify ¶
Whether the chain is internally consistent, optionally un-truncated and signed.
On its own this catches every edit, reorder, and deletion from the
middle of the chain: each of those leaves a record whose stored hash
no longer matches its contents, or whose prev_hash no longer points
at the record before it.
It cannot, on its own, catch a truncation. Dropping records from the
end — or discarding the trail entirely — leaves a shorter chain that is
perfectly valid on its own terms, so this returns True. That is a
property of hash chains in general, not of this implementation: nothing
inside a chain can attest to a link that was never handed to it.
Truncation is what expected_head is for. Persist :attr:head
somewhere the agent cannot reach — a WORM bucket, a append-only log, a
transparency log, a co-signer — and pass it back here. Every attack
above, truncation included, changes the head:
anchor = trail.head # written to durable, external storage
...
trail.verify(expected_head=anchor) # False if anything was removed
Nor can a chain alone catch a rebuild: rewrite a record and recompute
every hash after it, and the chain is consistent again. That is what
keys is for. With it, every record must carry a signature that
verifies under the public key its key_id names, so a rebuilt chain
fails unless it was signed with a key the verifier trusts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected_head
|
str | None
|
The chain head recorded out-of-band. When given, the trail must also end on this hash. Omit it and truncation goes undetected — see above. |
None
|
keys
|
Mapping[str, bytes | str] | None
|
Trusted public keys as PEM, by key id. When given, an unsigned record, an unknown key id, or a bad signature fails verification. |
None
|
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
was supplied, and every record is validly signed when |
Source code in .sdk/src/tulip/control/audit.py
export_jsonl ¶
The chain as newline-delimited JSON — one record per line, SIEM-ready.
from_records
classmethod
¶
Rebuild a trail from records (e.g. to :meth:verify an exported chain).
AuditRecord
dataclass
¶
AuditRecord(seq: int, ts: str, event_type: str, payload: dict[str, Any], prev_hash: str, hash: str, key_id: str | None = None, signature: str | None = None)
One link in the audit chain. hash commits to prev_hash.
AuditSigner ¶
Ed25519Signer ¶
An Ed25519 :class:AuditSigner.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
private_key
|
Any
|
A |
required |
key_id
|
str | None
|
The name verifiers look the public key up by. Defaults to the first 16 hex characters of the SHA-256 of the raw public key, so the same key always gets the same id. |
None
|
Source code in .sdk/src/tulip/control/audit.py
generate
classmethod
¶
A signer with a new random key.
from_pem
classmethod
¶
from_pem(pem: bytes | str, *, password: bytes | None = None, key_id: str | None = None) -> Ed25519Signer
A signer from a PEM-encoded (PKCS#8) Ed25519 private key.
Source code in .sdk/src/tulip/control/audit.py
private_key_pem ¶
The private key as PKCS#8 PEM, encrypted when password is given.
Source code in .sdk/src/tulip/control/audit.py
public_key_pem ¶
The public key as SubjectPublicKeyInfo PEM — what a verifier needs.
Source code in .sdk/src/tulip/control/audit.py
verify_jsonl ¶
verify_jsonl(text: str, *, keys: Mapping[str, bytes | str] | None = None, expected_head: str | None = None) -> bool
Verify an exported trail from its JSONL alone.
Needs no Tulip runtime state: an auditor with the export and the public keys runs this and nothing else. A line that is not a record fails.
Source code in .sdk/src/tulip/control/audit.py
AuditHook ¶
Bases: HookProvider
Records the agent's lifecycle into a tamper-evident :class:AuditTrail.
Source code in .sdk/src/tulip/control/governed.py
on_iteration_start
async
¶
Called at the start of each agent iteration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
iteration
|
int
|
Current iteration number (0-indexed) |
required |
state
|
AgentState
|
Current agent state |
required |
Source code in .sdk/src/tulip/hooks/provider.py
on_iteration_end
async
¶
Called at the end of each agent iteration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
iteration
|
int
|
Current iteration number (0-indexed) |
required |
state
|
AgentState
|
Current agent state |
required |
Source code in .sdk/src/tulip/hooks/provider.py
on_before_model_call
async
¶
Called before each model.complete() call.
Modify event.messages to change what the model sees. event.tools is read-only (inspect only).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
event
|
BeforeModelCallEvent
|
Write-protected event. Writable: messages. |
required |
Source code in .sdk/src/tulip/hooks/provider.py
on_after_model_call
async
¶
Called after each model.complete() call.
Set event.retry = True to discard response and re-call. Set event.response to replace the response. event.messages is read-only.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
event
|
AfterModelCallEvent
|
Write-protected event. Writable: response, retry. |
required |
Source code in .sdk/src/tulip/hooks/provider.py
Governed agents¶
An Agent pre-wired with grounding, guardrails, and an audit trail.
governed_agent ¶
governed_agent(model: Any = None, tools: list[Any] | None = None, *, system_prompt: str | None = None, profile: GovernanceProfile | None = None, audit_trail: AuditTrail | None = None, hooks: list[Any] | None = None, **kwargs: Any) -> GovernedAgent
Build a secure-by-default agent: grounded, guarded, and audited.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
Any
|
Model string or instance (as :class: |
None
|
tools
|
list[Any] | None
|
Tools available to the agent. |
None
|
system_prompt
|
str | None
|
System prompt. |
None
|
profile
|
GovernanceProfile | None
|
Which controls to enable (default: all on). |
None
|
audit_trail
|
AuditTrail | None
|
Reuse an existing trail; one is created if omitted. |
None
|
hooks
|
list[Any] | None
|
Extra hooks to add alongside the security hooks. |
None
|
**kwargs
|
Any
|
Passed through to :class: |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
A |
GovernedAgent
|
class: |
Source code in .sdk/src/tulip/control/governed.py
GovernedAgent
dataclass
¶
A secure-by-default :class:tulip.Agent plus its audit trail.
run / run_sync pass through to the wrapped agent; audit_trail
is the tamper-evident record of everything it did.
arun
async
¶
Async, thread-free twin of run_sync — delegates to the wrapped
agent's arun so a governed agent runs where threads aren't
available (e.g. the browser / Pyodide).
Source code in .sdk/src/tulip/control/governed.py
GovernanceProfile
dataclass
¶
Which secure-by-default controls a :func:governed_agent turns on.
All on by default — that is what makes the agent secure out of the box.
Verification¶
Evidence quality and adversarial refutation, feeding the
require_verification_score and min_severity rules on a policy.
verify
async
¶
verify(finding: FindingLike, *, skeptics: Sequence[Skeptic] | None = None, threshold: float = 0.6) -> VerificationResult
Independently challenge a finding; return whether it survives.
Runs each skeptic (default: a single :class:EvidenceQualitySkeptic),
collects their refutations, and re-grades confidence as the grounding score
minus the refutation penalties — where non-fatal penalties are capped
(:data:_MAX_NONFATAL_PENALTY) so volume of caveats alone can't refute a
well-grounded finding; a single fatal refutation zeroes it outright. A
finding survives only if nothing fatal was raised and confidence clears
threshold.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
finding
|
FindingLike
|
A :class: |
required |
skeptics
|
Sequence[Skeptic] | None
|
The challenge panel; defaults to the deterministic skeptic. Plug semantic/LLM skeptics here. |
None
|
threshold
|
float
|
Minimum confidence to survive (default 0.6). |
0.6
|
Returns:
| Name | Type | Description |
|---|---|---|
A |
VerificationResult
|
class: |
Source code in .sdk/src/tulip/control/verification.py
VerificationResult
dataclass
¶
VerificationResult(survives: bool, confidence: float, evidence_quality: float, refutations: list[Refutation] = list(), alternatives: list[str] = list(), notes: str = '')
The outcome of verifying a finding.
survives is False if any refutation is fatal or confidence falls
below the threshold. alternatives is populated by semantic skeptics
(the deterministic one leaves it empty).
Evidence ¶
Bases: BaseModel
A grounded security finding.
The gsar_score and evidence_refs fields are required: a
Evidence always knows how strongly it is grounded and what it is
grounded in. Build findings via :func:tulip.control.ground_finding
rather than constructing them directly — that is the path that
enforces the grounding threshold.
Severity ¶
Bases: StrEnum
Ordered severity band. StrEnum so it serialises as the bare string.
Not directly comparable with < (string ordering would be wrong);
use :func:severity_at_least or :data:SEVERITY_ORDER for ranking.