Skip to content

Control

The admission gate: decide whether a consequential action may run, and record the decision either way.

tulip.control is the domain-neutral surface, and the import path to use. Older releases shipped these under tulip.security; that path still resolves, with a deprecation warning, until 3.0.

For the concepts, start with The control layer and Writing a policy that holds.

Admitting an action

admit() evaluates the policy, records the decision on the audit trail, and runs the action only if it was allowed. A held or denied action raises AdmissionError carrying the ApprovalDecision that explains why.

admit async

admit(action: Action, perform: Callable[[], Awaitable[T]], *, policy: ControlPolicy, finding: Evidence | None = None, verdict: VerificationResult | None = None, trail: AuditTrail | None = None, approved_by: str | None = None, ledger: SpendLedger | None = None, spend_scope: str = 'default') -> T

Run perform only if action clears the trust chain; else reject.

The mandatory gate that turns the composable chain into an enforced one:

  1. :func:~tulip.control.policy.approve weighs the action against the evidence (finding), the verification (verdict), and the policy.
  2. The decision is recorded to trail (if given) — admitted or not — so no side effect is un-audited.
  3. On ALLOW, perform is awaited and its result returned. On require_human or deny, :class:AdmissionError is raised with the decision attached.

Parameters:

Name Type Description Default
action Action

The proposed side-effecting action.

required
perform Callable[[], Awaitable[T]]

A zero-arg async callable that performs the side effect.

required
policy ControlPolicy

The governing :class:~tulip.control.policy.ControlPolicy.

required
finding Evidence | None

The evidence the action responds to.

None
verdict VerificationResult | None

The :func:~tulip.control.verification.verify result.

None
trail AuditTrail | None

An :class:~tulip.control.audit.AuditTrail to record the decision on.

None
approved_by str | None

Who approved a require_human hold. With it, the held action runs and the approver is recorded on the trail. A deny still raises: no person approves past a denial.

None
ledger SpendLedger | None

Where spend_scope's cumulative spend is read before the decision and action.cost_usd recorded after perform succeeds.

None
spend_scope str

The scope the spend counts against: a customer, a tenant.

'default'

Returns:

Type Description
T

Whatever perform returns.

Raises:

Type Description
AdmissionError

if the action is not admitted (deny, or require_human without approved_by).

Source code in .sdk/src/tulip/control/admission.py
async def admit(
    action: Action,
    perform: Callable[[], Awaitable[T]],
    *,
    policy: ControlPolicy,
    finding: Evidence | None = None,
    verdict: VerificationResult | None = None,
    trail: AuditTrail | None = None,
    approved_by: str | None = None,
    ledger: SpendLedger | None = None,
    spend_scope: str = "default",
) -> T:
    """Run ``perform`` only if ``action`` clears the trust chain; else reject.

    The mandatory gate that turns the composable chain into an enforced one:

    1. :func:`~tulip.control.policy.approve` weighs the action against the evidence
       (``finding``), the verification (``verdict``), and the ``policy``.
    2. The decision is recorded to ``trail`` (if given) — admitted or not — so no
       side effect is un-audited.
    3. On ALLOW, ``perform`` is awaited and its result returned. On require_human or
       deny, :class:`AdmissionError` is raised with the decision attached.

    Args:
        action: The proposed side-effecting action.
        perform: A zero-arg async callable that performs the side effect.
        policy: The governing :class:`~tulip.control.policy.ControlPolicy`.
        finding: The evidence the action responds to.
        verdict: The :func:`~tulip.control.verification.verify` result.
        trail: An :class:`~tulip.control.audit.AuditTrail` to record the decision on.
        approved_by: Who approved a ``require_human`` hold. With it, the held action
            runs and the approver is recorded on the trail. A ``deny`` still raises:
            no person approves past a denial.
        ledger: Where ``spend_scope``'s cumulative spend is read before the
            decision and ``action.cost_usd`` recorded after ``perform`` succeeds.
        spend_scope: The scope the spend counts against: a customer, a tenant.

    Returns:
        Whatever ``perform`` returns.

    Raises:
        AdmissionError: if the action is not admitted (deny, or require_human
            without ``approved_by``).
    """
    spent = ledger.spent(spend_scope) if ledger is not None else 0.0
    decision = approve(action, policy=policy, finding=finding, verdict=verdict, spent_usd=spent)
    human = approved_by is not None and decision.outcome == ApprovalOutcome.REQUIRE_HUMAN
    if trail is not None:
        entry: dict[str, Any] = {
            "action": action.name,
            "asset": action.asset,
            "outcome": decision.outcome,
            "reason": decision.reason,
        }
        if ledger is not None:
            entry.update(
                {"cost_usd": action.cost_usd, "spent_usd": spent, "spend_scope": spend_scope}
            )
        if human:
            entry["approved_by"] = approved_by
        trail.record("action-admission", entry)
    if not (decision.allowed or human):
        raise AdmissionError(decision)
    result = await perform()
    if ledger is not None and action.cost_usd:
        # Only after it ran: a refused or failed action spends nothing.
        ledger.record(spend_scope, action.cost_usd, action=action.name)
    return result

AdmissionError

AdmissionError(decision: ApprovalDecision)

Bases: Exception

A side-effecting action failed admission — it did not clear the trust chain.

Carries the :class:~tulip.control.policy.ApprovalDecision so the caller can route a require_human hold to an approver or surface a deny reason.

Source code in .sdk/src/tulip/control/admission.py
def __init__(self, decision: ApprovalDecision) -> None:
    self.decision = decision
    super().__init__(
        f"action {decision.action.name!r} not admitted ({decision.outcome}): {decision.reason}"
    )

Deciding

approve() is the pure decision function — no I/O, no side effects. It takes an action and a policy and returns the outcome. Rules combine by taking the strongest result, so deny beats hold beats allow (in the API, ApprovalOutcome.DENY > REQUIRE_HUMAN > ALLOW).

approve

approve(action: Action, *, policy: ControlPolicy, finding: Evidence | None = None, verdict: VerificationResult | None = None, advisor: ControlAdvisor | None = None, spent_usd: float = 0.0) -> ApprovalDecision

Decide whether action may proceed: allow / require_human / deny.

Weighs every rule and returns the strongest triggered outcome (deny > require_human > allow), recording each check that fired so the decision is auditable.

Parameters:

Name Type Description Default
action Action

The proposed action.

required
policy ControlPolicy

The governing :class:ControlPolicy.

required
finding Evidence | None

The evidence the action responds to (optional).

None
verdict VerificationResult | None

The :func:~tulip.control.verification.verify result (optional, but auto-allow needs one that clears the policy bar).

None
spent_usd float

What the action's scope has already spent, for policy.spend_limit_usd.

0.0
advisor ControlAdvisor | None

An optional trained control model. It may only raise the decision toward caution — see :func:_combine. Omitting it, or passing one that fails, yields exactly the decision policy alone would have made.

None

Returns:

Name Type Description
An ApprovalDecision

class:ApprovalDecision.

Source code in .sdk/src/tulip/control/policy.py
def approve(
    action: Action,
    *,
    policy: ControlPolicy,
    finding: Evidence | None = None,
    verdict: VerificationResult | None = None,
    advisor: ControlAdvisor | None = None,
    spent_usd: float = 0.0,
) -> ApprovalDecision:
    """Decide whether ``action`` may proceed: allow / require_human / deny.

    Weighs every rule and returns the **strongest** triggered outcome (deny >
    require_human > allow), recording each check that fired so the decision is
    auditable.

    Args:
        action: The proposed action.
        policy: The governing :class:`ControlPolicy`.
        finding: The evidence the action responds to (optional).
        verdict: The :func:`~tulip.control.verification.verify` result (optional, but
            auto-allow needs one that clears the policy bar).
        spent_usd: What the action's scope has already spent, for
            ``policy.spend_limit_usd``.
        advisor: An optional trained control model. It may only raise the
            decision toward caution — see :func:`_combine`. Omitting it, or
            passing one that fails, yields exactly the decision policy alone
            would have made.

    Returns:
        An :class:`ApprovalDecision`.
    """
    triggered: list[tuple[str, str]] = []
    labels = action.labels()

    denied = labels & policy.deny_for
    if denied:
        triggered.append((ApprovalOutcome.DENY, f"labels {sorted(denied)} are denied by policy"))

    unsandboxed = labels & policy.require_sandbox_for
    if unsandboxed and SANDBOXED_TAG not in labels:
        triggered.append(
            (
                ApprovalOutcome.DENY,
                f"labels {sorted(unsandboxed)} require sandboxed execution "
                f"(the action carries no {SANDBOXED_TAG!r} tag)",
            )
        )

    if finding is not None and not severity_at_least(finding.severity, policy.min_severity):
        triggered.append(
            (
                ApprovalOutcome.DENY,
                f"finding severity {finding.severity.value} is below the policy "
                f"minimum {policy.min_severity.value}",
            )
        )

    if policy.require_verification_score > 0:
        if verdict is None:
            triggered.append((ApprovalOutcome.REQUIRE_HUMAN, "no verification provided"))
        elif not verdict.survives:
            triggered.append((ApprovalOutcome.DENY, "the finding did not survive verification"))
        elif verdict.confidence < policy.require_verification_score:
            triggered.append(
                (
                    ApprovalOutcome.REQUIRE_HUMAN,
                    f"verification confidence {verdict.confidence:.2f} is below the bar "
                    f"{policy.require_verification_score:.2f}",
                )
            )

    if action.blast_radius > policy.max_blast_radius:
        triggered.append(
            (
                ApprovalOutcome.REQUIRE_HUMAN,
                f"blast radius {action.blast_radius} exceeds the maximum {policy.max_blast_radius}",
            )
        )

    if (
        policy.require_human_over_usd is not None
        and action.cost_usd > policy.require_human_over_usd
    ):
        triggered.append(
            (
                ApprovalOutcome.REQUIRE_HUMAN,
                f"cost ${action.cost_usd:,.2f} exceeds the per-action limit "
                f"${policy.require_human_over_usd:,.2f}",
            )
        )

    if policy.spend_limit_usd is not None and spent_usd + action.cost_usd > policy.spend_limit_usd:
        triggered.append(
            (
                ApprovalOutcome.DENY,
                f"spend ${spent_usd:,.2f} + ${action.cost_usd:,.2f} would exceed the spend "
                f"limit ${policy.spend_limit_usd:,.2f}",
            )
        )

    needs_human = labels & policy.require_human_for
    if needs_human:
        triggered.append(
            (ApprovalOutcome.REQUIRE_HUMAN, f"labels {sorted(needs_human)} require human approval")
        )

    if not triggered:
        policy_outcome = ApprovalOutcome.ALLOW
        # Empty, not seeded with a passing note: ``checks`` records what *fired*,
        # and a clean allow fires nothing. That is the existing contract and the
        # combiner must not quietly change it.
        checks: list[str] = []
    else:
        policy_outcome = max((o for o, _ in triggered), key=lambda o: _ORDER[o])
        checks = [why for _, why in triggered]

    outcome, model_outcome, advisory = _combine(action, policy_outcome, advisor)
    if advisory is not None:
        checks = [*checks, advisory]

    return ApprovalDecision(
        outcome=outcome,
        reason="; ".join(checks) if checks else "all policy checks passed",
        action=action,
        checks=checks,
        policy_outcome=policy_outcome,
        model_outcome=model_outcome,
    )

ControlPolicy dataclass

ControlPolicy(require_verification_score: float = 0.8, max_blast_radius: int = 1, require_human_for: frozenset[str] = (lambda: frozenset({'production'}))(), deny_for: frozenset[str] = frozenset(), min_severity: Severity = Severity.LOW, require_sandbox_for: frozenset[str] = frozenset(), version: str = '', require_human_over_usd: float | None = None, spend_limit_usd: float | None = None)

The CISO knobs. Defaults are conservative — auto-allow only the safe path.

  • require_verification_score: minimum :class:VerificationResult confidence to auto-allow; below it (or with no verdict) a human is required.
  • max_blast_radius: most assets an action may affect to auto-allow.
  • require_human_for: action labels (environment / kind / tag) that always need a human (default: anything in production).
  • deny_for: labels that are hard-denied outright.
  • min_severity: don't act on findings below this band.
  • require_sandbox_for: labels whose actions must execute in a sandbox — an action matching one of these is denied unless it carries the :data:SANDBOXED_TAG tag. Enforced at the agent loop's tool seam by :class:~tulip.tools.sandbox.SandboxEnforcerHook.
  • version: a label for this policy. When set it is bound into approval ids, so a decision made under one version is never redeemed under another.
  • require_human_over_usd: an action whose cost_usd is above this needs a human.
  • spend_limit_usd: an action that would take its scope's cumulative spend past this is denied — no approval overrides it. The caller supplies what the scope has spent (spent_usd), typically from a spend ledger.

ApprovalDecision dataclass

ApprovalDecision(outcome: str, reason: str, action: Action, checks: list[str] = list(), policy_outcome: str = ApprovalOutcome.ALLOW, model_outcome: str | None = None)

The outcome of weighing an action against evidence, verification, and policy.

escalated_by_model property

escalated_by_model: bool

Whether a control model made this decision stricter than policy alone.

ApprovalOutcome

Outcome labels (kept simple/stable as plain strings).

Describing an action

A policy matches on what an action is — its environment, kind, blast radius, and tags — never on the name of the tool performing it.

Action dataclass

Action(name: str, asset: str = '', blast_radius: int = 1, environment: str = 'unknown', kind: str = '', tags: frozenset[str] = frozenset(), cost_usd: float = 0.0)

A proposed response action, with the attributes policy reasons over.

labels

labels() -> set[str]

The environment / kind / tags as one label set for policy matching.

Source code in .sdk/src/tulip/control/policy.py
def labels(self) -> set[str]:
    """The environment / kind / tags as one label set for policy matching."""
    return {self.environment, self.kind, *self.tags} - {""}

Gating a tool

gate_tool puts the gate in front of a tool the agent already has. The returned tool keeps the original's name, description and parameter schema, so the model sees no difference and nothing else in the agent changes — which is what makes the control structural rather than advisory. There is nothing to notice, so nothing to talk around.

from tulip.control import ControlPolicy, gate_tool

agent = Agent(model=model, tools=[
    lookup_order,                                     # read-only, ungated
    gate_tool(issue_refund, policy=ControlPolicy()),  # gated
])

A refusal comes back to the model as a readable result naming the outcome and the reason, so the agent can explain the hold rather than the run ending in a traceback. That is on_refusal="return", the default. Pass on_refusal="raise" for a caller that would rather stop, or on_refusal="interrupt" to pause the run on a hold until someone decides; a denial still comes back as a refusal (see Pausing until someone decides).

Gating a sandboxed tool composes rather than replacing it: the gate decides, and only an admitted call reaches the sandbox.

gate_tool

gate_tool(tool: Tool, *, policy: ControlPolicy, action: ActionSpec | None = None, trail: AuditTrail | None = None, finding: Evidence | None = None, verdict: VerificationResult | None = None, on_refusal: Literal['return', 'raise', 'interrupt'] = 'return', approval: ApprovalBridge | None = None, principal: str = 'agent', refusal_reason: str | Callable[[ApprovalDecision], str] | None = None, approval_context: Mapping[str, str] | Callable[[str, dict[str, Any]], Mapping[str, str]] | None = None, ledger: SpendLedger | None = None, spend_scope: str | Callable[[str, dict[str, Any]], str] = 'default') -> Tool

Return a copy of tool whose call goes through :func:admit first.

Parameters:

Name Type Description Default
tool Tool

The tool to gate. Not modified — an ungated reference stays usable, which matters when the same function is called by trusted code elsewhere.

required
policy ControlPolicy

The :class:ControlPolicy to weigh the call against.

required
action ActionSpec | None

How to turn a call into an :class:Action. A constant Action when risk does not vary, or (name, kwargs) -> Action when it does — the usual case, since a $12 refund and a $4,000,000 refund differ only in their arguments. None uses :func:~tulip.control.default_action, which tags the action with the tool's own name so a policy can still gate it by name.

None
trail AuditTrail | None

Records every decision, allowed or not. Omit and decisions are weighed but not written down.

None
finding Evidence | None

Grounded evidence supporting the action, when the policy requires one.

None
verdict VerificationResult | None

A verification result, when the policy sets require_verification_score.

None
approval ApprovalBridge | None

Where to submit an action held for a human. Without one, a hold tells the model it was held and stops there — true, and not actionable. With one, the refusal carries an approval_id the agent can poll. A denial never gets an id: it is final, and offering one would invite the agent to wait for a decision that is not coming.

None
principal str

Who the held action is attributed to on the approval.

'agent'
approval_context Mapping[str, str] | Callable[[str, dict[str, Any]], Mapping[str, str]] | None

What an approval is bound to besides the call itself, as a mapping or (tool_name, arguments) -> mapping: a thread, a tenant, a case id. It is part of the approval id, as is policy.version when set, so a decision made in one context or under one policy version is never redeemed in another.

None
ledger SpendLedger | None

A spend ledger. The gate reads the scope's cumulative spend before each decision, for policy.spend_limit_usd, and records action.cost_usd after the call runs.

None
spend_scope str | Callable[[str, dict[str, Any]], str]

The scope spend counts against, as a string or (tool_name, arguments) -> str: a customer, a tenant, a month.

'default'
on_refusal Literal['return', 'raise', 'interrupt']

"return" hands the model a JSON refusal naming the outcome and the reason, so it can explain itself to the user and the run continues. "raise" re-raises :class:AdmissionError for a caller that would rather stop. "interrupt" pauses the run on a require_human hold: the call returns the runtime's interrupt marker, the agent yields an InterruptEvent whose metadata carries the approval_id, and agent.resume(..., perform_dangling=True) re-issues the call once a person has decided on approval (which must be an :class:~tulip.control.ApprovalStore). Approved runs the call exactly once; denied returns a refusal. A policy deny never pauses.

'return'
refusal_reason str | Callable[[ApprovalDecision], str] | None

What the model is told when an action is refused. By default it is the policy's own reason, which names the checks that fired — accurate, and written in control-plane vocabulary the model will repeat verbatim to the end user ("blast radius 3 exceeds the maximum 1"). Pass a string, or (decision) -> str to vary by outcome, to say what the user should hear instead. The full policy reason is recorded on the audit trail regardless.

None

Returns:

Type Description
Tool

A new :class:~tulip.tools.decorator.Tool. Name, description and

Tool

parameter schema are the original's, so it is a drop-in wherever the

Tool

original was passed — the model cannot tell the difference, which is

Tool

the point: the gate is not something the model can be talked around.

"return" is the default because a refusal is information the agent can act on. An exception ends the run, and "the refund was held for a human" is something the user should hear rather than a stack trace.

Source code in .sdk/src/tulip/control/gate.py
def gate_tool(
    tool: Tool,
    *,
    policy: ControlPolicy,
    action: ActionSpec | None = None,
    trail: AuditTrail | None = None,
    finding: Evidence | None = None,
    verdict: VerificationResult | None = None,
    on_refusal: Literal["return", "raise", "interrupt"] = "return",
    approval: ApprovalBridge | None = None,
    principal: str = "agent",
    refusal_reason: str | Callable[[ApprovalDecision], str] | None = None,
    approval_context: (
        Mapping[str, str] | Callable[[str, dict[str, Any]], Mapping[str, str]] | None
    ) = None,
    ledger: SpendLedger | None = None,
    spend_scope: str | Callable[[str, dict[str, Any]], str] = "default",
) -> Tool:
    """Return a copy of ``tool`` whose call goes through :func:`admit` first.

    Args:
        tool: The tool to gate. Not modified — an ungated reference stays
            usable, which matters when the same function is called by trusted
            code elsewhere.
        policy: The :class:`ControlPolicy` to weigh the call against.
        action: How to turn a call into an :class:`Action`. A constant
            ``Action`` when risk does not vary, or ``(name, kwargs) -> Action``
            when it does — the usual case, since a $12 refund and a $4,000,000
            refund differ only in their arguments. ``None`` uses
            :func:`~tulip.control.default_action`, which tags the action with
            the tool's own name so a policy can still gate it by name.
        trail: Records every decision, allowed or not. Omit and decisions are
            weighed but not written down.
        finding: Grounded evidence supporting the action, when the policy
            requires one.
        verdict: A verification result, when the policy sets
            ``require_verification_score``.
        approval: Where to submit an action held for a human. Without one, a
            hold tells the model it was held and stops there — true, and not
            actionable. With one, the refusal carries an ``approval_id`` the
            agent can poll. A denial never gets an id: it is final, and
            offering one would invite the agent to wait for a decision that is
            not coming.
        principal: Who the held action is attributed to on the approval.
        approval_context: What an approval is bound to besides the call itself,
            as a mapping or ``(tool_name, arguments) -> mapping``: a thread, a
            tenant, a case id. It is part of the approval id, as is
            ``policy.version`` when set, so a decision made in one context or
            under one policy version is never redeemed in another.
        ledger: A spend ledger. The gate reads the scope's cumulative spend
            before each decision, for ``policy.spend_limit_usd``, and records
            ``action.cost_usd`` after the call runs.
        spend_scope: The scope spend counts against, as a string or
            ``(tool_name, arguments) -> str``: a customer, a tenant, a month.
        on_refusal: ``"return"`` hands the model a JSON refusal naming the
            outcome and the reason, so it can explain itself to the user and
            the run continues. ``"raise"`` re-raises
            :class:`AdmissionError` for a caller that would rather stop.
            ``"interrupt"`` pauses the run on a ``require_human`` hold: the call
            returns the runtime's interrupt marker, the agent yields an
            ``InterruptEvent`` whose ``metadata`` carries the ``approval_id``, and
            ``agent.resume(..., perform_dangling=True)`` re-issues the call once a
            person has decided on ``approval`` (which must be an
            :class:`~tulip.control.ApprovalStore`). Approved runs the call exactly
            once; denied returns a refusal. A policy ``deny`` never pauses.
        refusal_reason: What the model is told when an action is refused. By
            default it is the policy's own reason, which names the checks that
            fired — accurate, and written in control-plane vocabulary the model
            will repeat verbatim to the end user ("blast radius 3 exceeds the
            maximum 1"). Pass a string, or ``(decision) -> str`` to vary by
            outcome, to say what the user should hear instead. The full policy
            reason is recorded on the audit trail regardless.

    Returns:
        A new :class:`~tulip.tools.decorator.Tool`. Name, description and
        parameter schema are the original's, so it is a drop-in wherever the
        original was passed — the model cannot tell the difference, which is
        the point: the gate is not something the model can be talked around.

    ``"return"`` is the default because a refusal is information the agent can
    act on. An exception ends the run, and "the refund was held for a human" is
    something the user should hear rather than a stack trace.
    """
    from tulip.tools.decorator import Tool  # noqa: PLC0415 — avoids a cycle

    if on_refusal == "interrupt" and not isinstance(approval, ApprovalStore):
        raise TypeError(
            'on_refusal="interrupt" needs approval= an ApprovalStore (InMemoryApprovals, '
            "FileApprovals, or your own): a paused run has to find its decision somewhere"
        )

    inner = tool.fn
    # A sandboxed tool must keep running in its sandbox. `Tool.execute` returns
    # early for `sandbox is not None` and never reaches `fn`, so the two cannot
    # simply be stacked: carrying the sandbox onto the wrapper would skip the
    # gate, and dropping it -- as the first version of this did -- silently
    # moves the body back onto the host. Both fail quietly, which for a
    # security feature is the worst available outcome.
    #
    # They compose in one order only: gate first, then hand the admitted call
    # to the ORIGINAL tool, whose own `execute` still does the sandboxing.
    sandboxed = tool.sandbox is not None

    async def hold(
        error: AdmissionError,
        resolved: Any,
        kwargs: dict[str, Any],
        perform_with: Callable[[dict[str, Any]], Awaitable[Any]],
    ) -> Any:
        """A ``require_human`` hold in interrupt mode: pause, or act on a decision."""
        store = cast("ApprovalStore", approval)
        context = _approval_context(policy, approval_context, tool.name, kwargs)
        approval_id = store.submit(
            principal,
            tool.name,
            kwargs,
            reason=error.decision.reason,
            labels=sorted(resolved.labels()),
            context=context,
        )
        record = store.get(approval_id)
        where = {"approval_id": approval_id, "action": resolved.name, "asset": resolved.asset}

        if record is not None and record.status in ("approved", "denied"):
            if trail is not None:
                trail.record(
                    "approval-decision",
                    {
                        **where,
                        "verdict": record.status,
                        "decided_by": record.decided_by,
                        "approvers": record.approvers,
                    },
                )
            if record.status == "approved":
                edited = record.approved_arguments
                run_kwargs = dict(edited) if edited is not None else kwargs
                run_action = (
                    resolve_action(action, tool.name, run_kwargs)
                    if edited is not None
                    else resolved
                )
                if edited is not None and trail is not None:
                    trail.record(
                        "approval-edited",
                        {
                            **where,
                            "requested_arguments": dict(kwargs),
                            "approved_arguments": edited,
                        },
                    )

                async def perform_once() -> Any:
                    # Consumed before the side effect: a crash mid-call leaves
                    # the approval spent, never a second execution on replay.
                    store.consume(approval_id)
                    return await perform_with(run_kwargs)

                try:
                    # An edited call is weighed again: an approver cannot edit
                    # an action into one the policy denies.
                    return await admit(
                        run_action,
                        perform_once,
                        policy=policy,
                        finding=finding,
                        verdict=verdict,
                        trail=trail,
                        approved_by=record.decided_by,
                        ledger=ledger,
                        spend_scope=_scope_for(spend_scope, tool.name, run_kwargs),
                    )
                except AdmissionError as denial:
                    store.consume(approval_id)
                    return _refusal(denial, reason=refusal_reason)
            store.consume(approval_id)
            return json.dumps(
                {
                    "status": "denied",
                    "outcome": error.decision.outcome,
                    "action": resolved.name,
                    "asset": resolved.asset,
                    "reason": f"denied by {record.decided_by}",
                    "approval_id": approval_id,
                }
            )

        if trail is not None:
            trail.record("approval-requested", {**where, "principal": principal})
        return json.dumps(
            {
                "__interrupt__": True,
                "question": f"Approve {resolved.name} on {resolved.asset}? "
                + _reason_for(error, refusal_reason),
                "metadata": {
                    **where,
                    "principal": principal,
                    "reason": error.decision.reason,
                    "arguments": dict(kwargs),
                    "approvers": record.approvers if record is not None else [],
                    "context": context,
                },
            },
            default=str,
        )

    async def perform_with(call_kwargs: dict[str, Any]) -> Any:
        if sandboxed:
            return await tool.execute(**call_kwargs)
        result = inner(**call_kwargs)
        return await result if inspect.isawaitable(result) else result

    async def gated(**kwargs: Any) -> Any:
        async def perform() -> Any:
            return await perform_with(kwargs)

        resolved = resolve_action(action, tool.name, kwargs)
        try:
            return await admit(
                resolved,
                perform,
                policy=policy,
                finding=finding,
                verdict=verdict,
                trail=trail,
                ledger=ledger,
                spend_scope=_scope_for(spend_scope, tool.name, kwargs),
            )
        except AdmissionError as error:
            if on_refusal == "raise":
                raise
            if on_refusal == "interrupt" and error.decision.outcome != ApprovalOutcome.DENY:
                return await hold(error, resolved, kwargs, perform_with)
            return _refusal(
                error,
                approval=approval,
                principal=principal,
                kwargs=kwargs,
                reason=refusal_reason,
            )

    # `sandbox` is deliberately not set on the wrapper: it would short-circuit
    # `execute` and skip the gate. The sandbox is not lost — `perform` above
    # delegates to the original tool, which still has it.
    return Tool(
        name=tool.name,
        description=tool.description,
        parameters=tool.parameters,
        fn=gated,
        idempotent=tool.idempotent,
        labels=tool.labels,
    )

Holding an action for a human

A hold is only useful if the agent can find out what happened next. Give gate_tool an approval bridge and a held refusal carries an approval_id the agent can poll, while a human decides on a channel the agent cannot reach:

{"status": "held_for_approval", "outcome": "require_human",
 "action": "issue_refund", "asset": "ord-4821", "reason": "...",
 "approval_id": "appr-77",
 "next": "call approval_status(approval_id) once a human decides"}

A denial deliberately gets no id. It is final, and offering one would invite the agent to wait for a decision that is not coming.

ApprovalBridge is a structural Protocol with no import-time dependency, so any approval queue with submit and state methods satisfies it. The two stores Tulip ships, InMemoryApprovals and FileApprovals, satisfy it too; see Pausing until someone decides.

ApprovalBridge

Bases: Protocol

Submit a held action for out-of-band approval, and check its state.

A structural Protocol, deliberately: it has no import-time dependency on anything, so any approval queue with these two methods satisfies it.

Without one, a held action tells the model it was held and stops there — true, and not actionable. With one, the refusal carries an id the agent can poll while a human decides on a channel the agent cannot reach.

submit

submit(principal: str, tool: str, args: Mapping[str, Any]) -> str

Record a pending approval; return an id the agent can poll.

Source code in .sdk/src/tulip/control/gate.py
def submit(self, principal: str, tool: str, args: Mapping[str, Any]) -> str:
    """Record a pending approval; return an id the agent can poll."""
    ...

state

state(approval_id: str) -> str | None

Current state for an id — "pending" / "approved" / "denied".

Source code in .sdk/src/tulip/control/gate.py
def state(self, approval_id: str) -> str | None:
    """Current state for an id — ``"pending"`` / ``"approved"`` / ``"denied"``."""
    ...

What the user hears when an action is refused

The reason in that payload is, by default, the policy's own — a join of the checks that fired:

"blast radius 3 exceeds the maximum 1; labels ['large_refund'] are denied by policy"

That is the right level of detail for the audit trail and for a developer reading a log. It is also control-plane vocabulary, and a model handed it repeats it verbatim. Run against a live model, the refusal above reached the customer as "the blast radius (3) exceeds the maximum 1" and "it's classified as a large_refund".

refusal_reason gives the model the sentence you want the user to hear instead. Pass a string, or (decision) -> str to vary it by outcome:

gate_tool(
    issue_refund,
    policy=policy,
    trail=trail,
    refusal_reason=lambda d: (
        "We can't refund this amount automatically."
        if d.outcome == "deny"
        else "This refund is waiting on a manager."
    ),
)

The full policy reason still goes to the trail — a friendlier sentence for the customer must not shrink the record. Added in 2.10.0.

Pausing until someone decides

A polled id keeps the run going while a human decides. With on_refusal="interrupt" and an approval store, a hold pauses the run instead: the agent yields an InterruptEvent whose metadata carries the approval_id, the checkpointer keeps the conversation, and the store keeps the pending approval. A person decides with decide(), and agent.resume(..., thread_id=..., perform_dangling=True) re-issues the held call, which finds the decision:

from tulip.control import FileApprovals, gate_tool

store = FileApprovals("approvals.json")
refund = gate_tool(
    issue_refund,
    policy=policy,
    approval=store,
    on_refusal="interrupt",
    trail=trail,
)
agent = Agent(model=model, tools=[refund], checkpointer=FileCheckpointer("checkpoints"))

async for event in agent.run("refund order 4821", thread_id="t1"):
    if isinstance(event, InterruptEvent):
        approval_id = event.metadata["approval_id"]   # the run is parked

# Later, from any process that can open the same file:
store.decide(approval_id, "approved", by="[email protected]")
async for event in agent.resume("approved", thread_id="t1", perform_dangling=True):
    ...

An approval names one call. Its id is derived from a digest of the principal, the tool, the arguments, policy.version when set, and any approval_context, so a call with different arguments waits for its own decision. An approved call is weighed against the policy again when it is redeemed, so a deny such as a spend limit still refuses it, and it runs at most once as long as the store's writes do not race (see the FileApprovals limit below): the approval is consumed immediately before the side effect, so a repeated call holds again instead of riding an old yes. A denied call returns a refusal. A policy deny never pauses. on_refusal="interrupt" needs an ApprovalStore, not a bare ApprovalBridge, and raises TypeError without one.

InMemoryApprovals lives in one process and is gone on restart; it is for tests and demos. FileApprovals keeps every record in one JSON file, re-read on every call and written atomically, so a decision written by another process is seen by the next call. Writes are serialised only through one FileApprovals object in one process, and the agent's own submits and consumes are writes too. A decision saved from another process, or through another FileApprovals object, at the same moment can overwrite one of them. In the worst case a consumed approval goes back to approved and can be redeemed again. Where that matters, keep every writer in one process and give them the same FileApprovals object, or implement ApprovalStore over a database with atomic updates.

Who may decide. Without an ApprovalAuthority, any named principal can decide. With one, every decision is checked when it is made:

from tulip.control import ApprovalAuthority, ApproverRule, FileApprovals

authority = ApprovalAuthority(
    rules=(
        ApproverRule(labels=frozenset({"payment"}), roles=frozenset({"finance"})),
        ApproverRule(
            labels=frozenset({"production"}),
            approvers=frozenset({"olga", "sam"}),
            quorum=2,
        ),
    ),
    roles_of=directory.roles_for,
)
store = FileApprovals("approvals.json", authority=authority)

Every rule that matches the held action's labels must reach its quorum of distinct authorised approvers before the call is approved; one authorised denial ends it. The principal that requested the action cannot approve it, directly or through a delegation it granted, unless every matching rule sets allow_self_approval. A Delegation lends an approver's authority to someone else until a deadline. A decision by someone without authority raises ApprovalAuthorityError and is kept on the record as a rejection. An action no rule covers cannot be approved by anyone: the authority fails closed.

ApprovalStore

Bases: Protocol

Where held calls wait for a decision.

A superset of :class:~tulip.control.ApprovalBridge: submit and state keep that shape, so a store works anywhere a bridge does.

submit

submit(principal: str, tool: str, args: Mapping[str, Any], *, reason: str = '', labels: Iterable[str] = (), context: Mapping[str, str] | None = None) -> str

Record a held call, or return the live record for the same call.

Source code in .sdk/src/tulip/control/approvals.py
def submit(
    self,
    principal: str,
    tool: str,
    args: Mapping[str, Any],
    *,
    reason: str = "",
    labels: Iterable[str] = (),
    context: Mapping[str, str] | None = None,
) -> str:
    """Record a held call, or return the live record for the same call."""
    ...

state

state(approval_id: str) -> str | None

The record's status, or None for an unknown id.

Source code in .sdk/src/tulip/control/approvals.py
def state(self, approval_id: str) -> str | None:
    """The record's status, or ``None`` for an unknown id."""
    ...

get

get(approval_id: str) -> ApprovalRecord | None

The full record, or None for an unknown id.

Source code in .sdk/src/tulip/control/approvals.py
def get(self, approval_id: str) -> ApprovalRecord | None:
    """The full record, or ``None`` for an unknown id."""
    ...

decide

decide(approval_id: str, verdict: Verdict, *, by: str, arguments: Mapping[str, Any] | None = None) -> ApprovalRecord

Approve or deny a pending record, naming who decided.

Source code in .sdk/src/tulip/control/approvals.py
def decide(
    self,
    approval_id: str,
    verdict: Verdict,
    *,
    by: str,
    arguments: Mapping[str, Any] | None = None,
) -> ApprovalRecord:
    """Approve or deny a pending record, naming who decided."""
    ...

consume

consume(approval_id: str) -> ApprovalRecord

Mark a decided record as acted on, so it cannot be used again.

Source code in .sdk/src/tulip/control/approvals.py
def consume(self, approval_id: str) -> ApprovalRecord:
    """Mark a decided record as acted on, so it cannot be used again."""
    ...

pending

pending() -> list[ApprovalRecord]

Records still waiting for a decision, oldest first.

Source code in .sdk/src/tulip/control/approvals.py
def pending(self) -> list[ApprovalRecord]:
    """Records still waiting for a decision, oldest first."""
    ...

InMemoryApprovals

InMemoryApprovals(authority: ApprovalAuthority | None = None)

Bases: _Approvals

An approval store in this process only. Gone on restart; for tests and demos.

Source code in .sdk/src/tulip/control/approvals.py
def __init__(self, authority: ApprovalAuthority | None = None) -> None:
    super().__init__(authority)
    self._records: dict[str, ApprovalRecord] = {}

FileApprovals

FileApprovals(path: str | Path, authority: ApprovalAuthority | None = None)

Bases: _Approvals

An approval store in one JSON file.

Every call re-reads the file, so a decision written by another process is seen by the next call. Writes are atomic (temporary file, then rename), so a crash never leaves a half-written store. Concurrent writers are serialised within a process only: use one process to decide, or a database-backed store for many.

Source code in .sdk/src/tulip/control/approvals.py
def __init__(self, path: str | Path, authority: ApprovalAuthority | None = None) -> None:
    super().__init__(authority)
    self.path = Path(path)

ApprovalRecord dataclass

ApprovalRecord(approval_id: str, digest: str, principal: str, tool: str, arguments: dict[str, Any], status: Status = 'pending', reason: str = '', created_at: str = _now(), verdict: Verdict | None = None, decided_by: str | None = None, decided_at: str | None = None, labels: list[str] = list(), approvals: list[dict[str, Any]] = list(), rejections: list[dict[str, Any]] = list(), context: dict[str, str] = dict(), approved_arguments: dict[str, Any] | None = None)

One held call and what became of it.

approvers property

approvers: list[str]

Distinct principals whose approval was accepted, in order.

call_digest

call_digest(principal: str, tool: str, arguments: Mapping[str, Any], context: Mapping[str, str] | None = None) -> str

SHA-256 over the canonical form of one call.

Key order does not matter; any change to a value, the tool, the principal or the context does. context binds the call to where it was made, such as a policy version or a thread. An empty context hashes exactly like no context, so records written before contexts existed still match.

Source code in .sdk/src/tulip/control/approvals.py
def call_digest(
    principal: str,
    tool: str,
    arguments: Mapping[str, Any],
    context: Mapping[str, str] | None = None,
) -> str:
    """SHA-256 over the canonical form of one call.

    Key order does not matter; any change to a value, the tool, the principal
    or the context does. ``context`` binds the call to where it was made, such
    as a policy version or a thread. An empty context hashes exactly like no
    context, so records written before contexts existed still match.
    """
    body: dict[str, Any] = {"principal": principal, "tool": tool, "arguments": dict(arguments)}
    if context:
        body["context"] = dict(context)
    canonical = json.dumps(body, sort_keys=True, separators=(",", ":"), default=str)
    return hashlib.sha256(canonical.encode("utf-8")).hexdigest()

ApprovalAuthority dataclass

ApprovalAuthority(rules: tuple[ApproverRule, ...], delegations: tuple[Delegation, ...] = (), roles_of: Callable[[str], Iterable[str]] | None = None, clock: Callable[[], datetime] = _utcnow)

Approver rules, delegations, and how to look up a principal's roles.

Rules are evaluated by position, so keep their order stable for records that are still pending.

matching

matching(labels: Iterable[str]) -> dict[int, ApproverRule]

The rules governing an action carrying labels, by position.

Source code in .sdk/src/tulip/control/approvals.py
def matching(self, labels: Iterable[str]) -> dict[int, ApproverRule]:
    """The rules governing an action carrying ``labels``, by position."""
    wanted = frozenset(labels)
    return {i: rule for i, rule in enumerate(self.rules) if rule.matches(wanted)}

check

check(record: ApprovalRecord, by: str) -> _Basis

Which matching rules by may decide under, and on what basis.

Raises:

Type Description
ApprovalAuthorityError

No rule covers the action, by requested it, or by holds none of the matching rules, directly or by an active delegation.

Source code in .sdk/src/tulip/control/approvals.py
def check(self, record: ApprovalRecord, by: str) -> _Basis:
    """Which matching rules ``by`` may decide under, and on what basis.

    Raises:
        ApprovalAuthorityError: No rule covers the action, ``by`` requested
            it, or ``by`` holds none of the matching rules, directly or by an
            active delegation.
    """
    labels = frozenset(record.labels)
    rules = self.matching(labels)
    if not rules:
        raise ApprovalAuthorityError(
            f"no approver rule covers labels {sorted(labels)}; nobody may decide it"
        )
    self_allowed = all(rule.allow_self_approval for rule in rules.values())
    if by == record.principal and not self_allowed:
        raise ApprovalAuthorityError(f"{by} requested this action and may not decide it")

    direct = frozenset(i for i, rule in rules.items() if self._holds(rule, by))
    if direct:
        return _Basis(direct, by)

    now = self.clock()
    for grant in self.delegations:
        if grant.grantee != by or not grant.active(now):
            continue
        if grant.labels and not (grant.labels & labels):
            continue
        if grant.grantor == record.principal and not self_allowed:
            continue
        lent = frozenset(i for i, rule in rules.items() if self._holds(rule, grant.grantor))
        if lent:
            return _Basis(lent, f"delegated by {grant.grantor}")
    raise ApprovalAuthorityError(f"{by} may not decide {record.tool} (labels {sorted(labels)})")

satisfied

satisfied(record: ApprovalRecord, approvals: list[dict[str, Any]]) -> bool

Whether every matching rule has its quorum of distinct approvers.

Source code in .sdk/src/tulip/control/approvals.py
def satisfied(self, record: ApprovalRecord, approvals: list[dict[str, Any]]) -> bool:
    """Whether every matching rule has its quorum of distinct approvers."""
    for i, rule in self.matching(record.labels).items():
        approvers = {
            a["by"]
            for a in approvals
            if a.get("verdict") == "approved" and i in a.get("rules", ())
        }
        if len(approvers) < rule.quorum:
            return False
    return True

ApproverRule dataclass

ApproverRule(labels: frozenset[str] = frozenset(), approvers: frozenset[str] = frozenset(), roles: frozenset[str] = frozenset(), quorum: int = 1, allow_self_approval: bool = False)

Who may decide actions that carry some labels.

Parameters:

Name Type Description Default
labels frozenset[str]

Labels this rule governs. Empty matches every action.

frozenset()
approvers frozenset[str]

Principals allowed to decide, by name.

frozenset()
roles frozenset[str]

Roles allowed to decide, resolved per principal by :attr:ApprovalAuthority.roles_of.

frozenset()
quorum int

Distinct authorised approvals needed.

1
allow_self_approval bool

Whether the principal that requested the action may approve it. Off by default: separation of duties.

False

matches

matches(labels: frozenset[str]) -> bool

Whether this rule governs an action carrying labels.

Source code in .sdk/src/tulip/control/approvals.py
def matches(self, labels: frozenset[str]) -> bool:
    """Whether this rule governs an action carrying ``labels``."""
    return not self.labels or bool(self.labels & labels)

Delegation dataclass

Delegation(grantor: str, grantee: str, expires_at: str, labels: frozenset[str] = frozenset())

An approver lends their authority to someone else until a deadline.

Parameters:

Name Type Description Default
grantor str

The principal whose authority is lent.

required
grantee str

The principal who may use it.

required
expires_at str

ISO-8601 deadline; a naive timestamp is read as UTC.

required
labels frozenset[str]

Limit the delegation to actions carrying these labels. Empty lends everything the grantor may decide.

frozenset()

active

active(now: datetime) -> bool

Whether the deadline is still in the future.

Source code in .sdk/src/tulip/control/approvals.py
def active(self, now: datetime) -> bool:
    """Whether the deadline is still in the future."""
    deadline = datetime.fromisoformat(self.expires_at)
    if deadline.tzinfo is None:
        deadline = deadline.replace(tzinfo=UTC)
    return now < deadline

ApprovalAuthorityError

Bases: PermissionError

A decision was refused because the decider lacks authority for it.

Capping spend

Action.cost_usd says what an action spends. ControlPolicy can hold an action above a per-action cost (require_human_over_usd) and deny one that would take its scope's cumulative spend past a limit (spend_limit_usd); no approval overrides that denial. A spend ledger supplies the cumulative figure:

from tulip.control import Action, ControlPolicy, FileSpendLedger, gate_tool

refund = gate_tool(
    issue_refund,
    policy=ControlPolicy(spend_limit_usd=5_000, require_human_over_usd=500),
    action=lambda name, args: Action(
        name=name, asset=args["order_id"], cost_usd=args["amount_usd"]
    ),
    ledger=FileSpendLedger("spend.json"),
    spend_scope=lambda name, args: f"customer:{args['customer_id']}",
)

A scope is any string: a thread, a customer, a tenant, a month. Spend is recorded only after the action has run, so a refused or failed action costs nothing. The check and the record are not one transaction. Two calls against the same scope that overlap can each pass the check before either records, so the cap can be exceeded. That includes two processes, and also two tool calls the agent runs together in one turn (the default is tool_execution="concurrent"). SpendLedger reads and records in separate calls, before and after the action, so a ledger cannot make the check atomic by itself. Where the cap must hold, run the agent that has the spending tools with tool_execution="sequential" and keep one writer per scope. InMemorySpendLedger is the in-process version, for tests and demos. admit() takes ledger= and a string spend_scope= directly.

SpendLedger

Bases: Protocol

Cumulative spend per scope.

spent

spent(scope: str) -> float

USD recorded against scope so far; 0 for an unknown scope.

Source code in .sdk/src/tulip/control/spend.py
def spent(self, scope: str) -> float:
    """USD recorded against ``scope`` so far; 0 for an unknown scope."""
    ...

record

record(scope: str, amount_usd: float, *, action: str = '') -> float

Add amount_usd to scope and return the new total.

Source code in .sdk/src/tulip/control/spend.py
def record(self, scope: str, amount_usd: float, *, action: str = "") -> float:
    """Add ``amount_usd`` to ``scope`` and return the new total."""
    ...

InMemorySpendLedger

InMemorySpendLedger()

A spend ledger in this process only. Gone on restart; for tests and demos.

Source code in .sdk/src/tulip/control/spend.py
def __init__(self) -> None:
    self._lock = threading.Lock()
    self._totals: dict[str, float] = {}

FileSpendLedger

FileSpendLedger(path: str | Path)

A spend ledger in one JSON file, with every entry kept.

Re-read on every call and written atomically (temporary file, then rename), like :class:~tulip.control.FileApprovals. Writers are serialised within a process only.

Source code in .sdk/src/tulip/control/spend.py
def __init__(self, path: str | Path) -> None:
    self.path = Path(path)
    self._lock = threading.Lock()

entries

entries(scope: str) -> list[dict[str, Any]]

Every recorded spend for scope, oldest first.

Source code in .sdk/src/tulip/control/spend.py
def entries(self, scope: str) -> list[dict[str, Any]]:
    """Every recorded spend for ``scope``, oldest first."""
    with self._lock:
        return list(self._load().get(scope, {}).get("entries", []))

Deriving action labels

Turn a tool call into an Action using declarative rules, so the labels a policy matches on are not hand-written per call site.

ActionSpec module-attribute

ActionSpec = Action | Callable[[str, Mapping[str, Any]], Action]

resolve_action

resolve_action(spec: ActionSpec | None, name: str, kwargs: Mapping[str, Any]) -> Action

Resolve an :class:ActionSpec (or None) into a concrete :class:Action.

Source code in .sdk/src/tulip/control/action.py
def resolve_action(spec: ActionSpec | None, name: str, kwargs: Mapping[str, Any]) -> Action:
    """Resolve an :class:`ActionSpec` (or ``None``) into a concrete :class:`Action`."""
    if spec is None:
        return default_action(name, kwargs)
    if isinstance(spec, Action):
        return spec
    return spec(name, kwargs)

default_action

default_action(name: str, kwargs: Mapping[str, Any], *, environment: str = 'unknown', kind: str = '', blast_radius: int = 1, tags: frozenset[str] | None = None) -> Action

A conservative :class:Action for name when none was supplied.

Fail-safe by construction: environment="unknown" plus the stock :class:~tulip.control.policy.ControlPolicy (which requires a verification score) lands an un-verified call on require_human rather than auto-allowing it.

tags defaults to the action's own name, so a policy can always gate one specific tool by naming it — the one thing that worked before labels were derived at all.

Source code in .sdk/src/tulip/control/action.py
def default_action(
    name: str,
    kwargs: Mapping[str, Any],
    *,
    environment: str = "unknown",
    kind: str = "",
    blast_radius: int = 1,
    tags: frozenset[str] | None = None,
) -> Action:
    """A conservative :class:`Action` for ``name`` when none was supplied.

    Fail-safe by construction: ``environment="unknown"`` plus the stock
    :class:`~tulip.control.policy.ControlPolicy` (which requires a verification
    score) lands an un-verified call on ``require_human`` rather than
    auto-allowing it.

    ``tags`` defaults to the action's own name, so a policy can always gate one
    specific tool by naming it — the one thing that worked before labels were
    derived at all.
    """
    return Action(
        name=name,
        asset=asset_from_args(kwargs),
        blast_radius=blast_radius,
        environment=environment,
        kind=kind,
        tags=tags if tags is not None else frozenset({name}),
    )

action_from_labels

action_from_labels(name: str, kwargs: Mapping[str, Any], *, labels: Mapping[str, Any] | None = None, environment: str | None = None, blast_radius: int = 1) -> Action

Build an :class:Action from a tool's declared labels.

labels is what a tool definition declares about the actions it performs — environment, kind, blast_radius, tags. Anything absent falls back: the caller's environment (the agent's, or the deployment's), then "unknown".

The tool's own name is always among the tags, so naming a tool in require_human_for keeps working regardless of what it declares.

labels["derive"] may carry argument-derived rules (see :func:derive_labels) — the only part of this that reads kwargs for labelling. Derived tags join the declared ones, a derived set_kind / set_environment wins over the declared value (it describes this call), and a derived blast radius only ever raises the declared one. With no derive key the result is exactly what it was before.

Source code in .sdk/src/tulip/control/action.py
def action_from_labels(
    name: str,
    kwargs: Mapping[str, Any],
    *,
    labels: Mapping[str, Any] | None = None,
    environment: str | None = None,
    blast_radius: int = 1,
) -> Action:
    """Build an :class:`Action` from a tool's declared labels.

    ``labels`` is what a tool definition declares about the actions it performs
    — ``environment``, ``kind``, ``blast_radius``, ``tags``. Anything absent
    falls back: the caller's ``environment`` (the agent's, or the deployment's),
    then ``"unknown"``.

    The tool's own name is always among the tags, so naming a tool in
    ``require_human_for`` keeps working regardless of what it declares.

    ``labels["derive"]`` may carry argument-derived rules (see
    :func:`derive_labels`) — the only part of this that reads ``kwargs`` for
    labelling. Derived tags join the declared ones, a derived ``set_kind`` /
    ``set_environment`` wins over the declared value (it describes *this* call),
    and a derived blast radius only ever raises the declared one. With no
    ``derive`` key the result is exactly what it was before.
    """
    declared = dict(labels or {})
    declared_env = str(declared.get("environment") or "").strip()
    declared_kind = str(declared.get("kind") or "").strip()
    declared_tags = declared.get("tags") or []
    if isinstance(declared_tags, str):
        declared_tags = [declared_tags]

    radius = declared.get("blast_radius")
    try:
        resolved_radius = int(radius) if radius is not None else blast_radius
    except (TypeError, ValueError):
        resolved_radius = blast_radius

    derived = derive_labels(declared.get("derive"), kwargs)
    if derived.blast_radius is not None:
        resolved_radius = max(resolved_radius, derived.blast_radius)

    tags = {name, *(str(t) for t in declared_tags if t), *derived.tags}
    if derived.undetermined:
        tags.add(UNDETERMINED_TAG)

    return default_action(
        name,
        kwargs,
        environment=(
            derived.environment or declared_env or (environment or "").strip() or "unknown"
        ),
        kind=derived.kind or declared_kind,
        blast_radius=resolved_radius,
        tags=frozenset(tags),
    )

derive_labels

derive_labels(rules: Any, kwargs: Mapping[str, Any]) -> DerivedLabels

Evaluate a tool's derive rules against one call's arguments.

Declarative and total: comparisons only, never eval, never a callable, so nothing a tool receives can execute during labelling. Rules apply in order and all matching rules apply. Anything that cannot be evaluated — a missing argument, an argument of the wrong type, a malformed rule — is skipped and records :data:UNDETERMINED_TAG, so "we could not tell" reaches the policy as a fact rather than as silence.

Source code in .sdk/src/tulip/control/action.py
def derive_labels(rules: Any, kwargs: Mapping[str, Any]) -> DerivedLabels:
    """Evaluate a tool's ``derive`` rules against one call's arguments.

    Declarative and total: comparisons only, never ``eval``, never a callable,
    so nothing a tool receives can execute during labelling. Rules apply in
    order and all matching rules apply. Anything that cannot be evaluated — a
    missing argument, an argument of the wrong type, a malformed rule — is
    skipped *and* records :data:`UNDETERMINED_TAG`, so "we could not tell"
    reaches the policy as a fact rather than as silence.
    """
    derived = DerivedLabels()
    if rules is None:
        return derived
    if not isinstance(rules, list | tuple):
        derived.mark_undetermined()
        return derived
    for rule in rules:
        _apply_rule(rule, kwargs, derived)
    return derived

DerivedLabels

DerivedLabels()

Accumulator for what a derive list adds to an action.

Source code in .sdk/src/tulip/control/action.py
def __init__(self) -> None:
    self.tags: set[str] = set()
    self.kind: str = ""
    self.environment: str = ""
    self.blast_radius: int | None = None
    self.undetermined: bool = False

raise_radius

raise_radius(value: int) -> None

Deriving may raise the blast radius; it may never lower it.

Source code in .sdk/src/tulip/control/action.py
def raise_radius(self, value: int) -> None:
    """Deriving may raise the blast radius; it may never lower it."""
    self.blast_radius = value if self.blast_radius is None else max(self.blast_radius, value)

asset_from_args

asset_from_args(kwargs: Mapping[str, Any]) -> str

Best-effort asset label from a tool call's arguments.

Source code in .sdk/src/tulip/control/action.py
def asset_from_args(kwargs: Mapping[str, Any]) -> str:
    """Best-effort asset label from a tool call's arguments."""
    for key in _ASSET_KEYS:
        value = kwargs.get(key)
        if value:
            return str(value)
    return ""

UNDETERMINED_TAG module-attribute

UNDETERMINED_TAG = 'undetermined'

SANDBOXED_TAG module-attribute

SANDBOXED_TAG = 'sandboxed'

The record

A hash-chained log of every decision. Each record commits to the previous hash, so editing any record breaks verify().

Tamper-evident, not tamper-proof

By default this is a keyless SHA-256 chain held in memory. It detects edits when checked against a head hash you retain out-of-band; it does not prevent them, and it does not anchor the log. Without that head, anyone who can write the log can rebuild an unsigned chain around an edit. Persist the JSONL and pin the head hash externally before relying on it as compliance evidence, and sign the trail to catch a rebuild.

What verify() catches, and the one thing it cannot

Called with no arguments, verify() catches every edit, every reorder, and every deletion from the middle of the chain — each leaves a record whose stored hash no longer matches its contents, or whose prev_hash no longer points at the record before it.

It cannot, on its own, catch a truncation:

Attack verify() verify(expected_head=…)
Edit a record False False
Reorder records False False
Delete from the middle False False
Drop records off the end True False
Discard the trail entirely True False

Dropping the tail leaves a shorter chain that is perfectly valid on its own terms. That is a property of hash chains in general, not of this implementation: nothing inside a chain can attest to a link that was never handed to it. An agent that can reach its own audit trail can therefore erase the end of it and still pass a bare verify().

Anchoring closes it. Every attack in that table moves the head, so keep head somewhere the agent cannot reach — a WORM bucket, an append-only log, a co-signer, a transparency log — and pass it back:

trail = AuditTrail()
...
anchor = trail.head                      # to durable, external storage

# later, on the exported chain
restored = AuditTrail.from_records(records)
restored.verify(expected_head=anchor)    # False if anything was removed

Added in 2.10.0, alongside a correction: verify() previously documented itself as detecting "no edit, deletion, or reorder", which overstated what a chain can prove about its own tail.

Signing the trail

A hash chain proves the records agree with each other, not who wrote them. Rewrite a record and recompute every hash after it, and a bare verify() passes again. Give the trail a signer and every record's hash is signed with Ed25519, and the record carries the key_id of the key that signed it:

from tulip.control import AuditTrail, Ed25519Signer, verify_jsonl

trail = AuditTrail(signer=Ed25519Signer.from_pem(private_pem, key_id="audit-2026-09"))
...
anchor = trail.head                      # to durable, external storage
exported = trail.export_jsonl()

# An auditor with the export and the public key, and no Tulip runtime state:
verify_jsonl(exported, keys={"audit-2026-09": public_pem}, expected_head=anchor)

With keys, an unsigned record, an unknown key id, or a bad signature fails verification, so a chain rebuilt around an edit fails unless it was signed with a key the verifier trusts. trail.verify(keys=...) runs the same check in process.

Signing does not catch truncation. Records dropped off the end take their signatures with them, and what is left is still validly signed, so keep passing an externally held expected_head.

To rotate keys, call trail.use_signer(new_signer). Records already written keep the key they were signed with, so give the verifier both public keys. Signing needs the cryptography package (pip install "tulip-agents[audit]"); an unsigned trail needs nothing and exports exactly as before.

AuditTrail

AuditTrail(*, clock: Callable[[], str] | None = None, signer: AuditSigner | None = None)

An append-only, hash-chained log of agent actions.

Append with :meth:record (or :meth:record_event for a Tulip event); check integrity with :meth:verify; ship with :meth:export_jsonl. Pass clock to make timestamps deterministic in tests, and signer to sign every record.

Source code in .sdk/src/tulip/control/audit.py
def __init__(
    self,
    *,
    clock: Callable[[], str] | None = None,
    signer: AuditSigner | None = None,
) -> None:
    self._records: list[AuditRecord] = []
    self._clock = clock or _utc_now_iso
    self._signer = signer

head property

head: str

Hash of the latest record, or the genesis anchor when empty.

use_signer

use_signer(signer: AuditSigner | None) -> None

Sign records from now on with signer — how a key is rotated.

Records already written keep the key they were signed with; a verifier given both public keys accepts the whole trail.

Source code in .sdk/src/tulip/control/audit.py
def use_signer(self, signer: AuditSigner | None) -> None:
    """Sign records from now on with ``signer`` — how a key is rotated.

    Records already written keep the key they were signed with; a verifier
    given both public keys accepts the whole trail.
    """
    self._signer = signer

record

record(event_type: str, payload: Mapping[str, Any] | None = None) -> AuditRecord

Append a record committing to the current chain head.

Source code in .sdk/src/tulip/control/audit.py
def record(self, event_type: str, payload: Mapping[str, Any] | None = None) -> AuditRecord:
    """Append a record committing to the current chain head."""
    seq = len(self._records)
    prev = self.head
    ts = self._clock()
    body = dict(payload or {})
    digest = _entry_hash(seq, ts, event_type, body, prev)
    key_id = signature = None
    if self._signer is not None:
        key_id = self._signer.key_id
        signature = base64.b64encode(self._signer.sign(digest.encode("ascii"))).decode("ascii")
    rec = AuditRecord(
        seq=seq,
        ts=ts,
        event_type=event_type,
        payload=body,
        prev_hash=prev,
        hash=digest,
        key_id=key_id,
        signature=signature,
    )
    self._records.append(rec)
    return rec

record_event

record_event(event: Any) -> AuditRecord

Append a record for a Tulip event (duck-typed; safe scalar fields).

Source code in .sdk/src/tulip/control/audit.py
def record_event(self, event: Any) -> AuditRecord:
    """Append a record for a Tulip event (duck-typed; safe scalar fields)."""
    payload: dict[str, Any] = {}
    for key in ("name", "tool", "final_message", "reason", "content", "asset"):
        val = getattr(event, key, None)
        if isinstance(val, str | int | float | bool):
            payload[key] = val
    return self.record(type(event).__name__, payload)

records

records() -> list[AuditRecord]

A copy of the records, in order.

Source code in .sdk/src/tulip/control/audit.py
def records(self) -> list[AuditRecord]:
    """A copy of the records, in order."""
    return list(self._records)

verify

verify(*, expected_head: str | None = None, keys: Mapping[str, bytes | str] | None = None) -> bool

Whether the chain is internally consistent, optionally un-truncated and signed.

On its own this catches every edit, reorder, and deletion from the middle of the chain: each of those leaves a record whose stored hash no longer matches its contents, or whose prev_hash no longer points at the record before it.

It cannot, on its own, catch a truncation. Dropping records from the end — or discarding the trail entirely — leaves a shorter chain that is perfectly valid on its own terms, so this returns True. That is a property of hash chains in general, not of this implementation: nothing inside a chain can attest to a link that was never handed to it.

Truncation is what expected_head is for. Persist :attr:head somewhere the agent cannot reach — a WORM bucket, a append-only log, a transparency log, a co-signer — and pass it back here. Every attack above, truncation included, changes the head:

anchor = trail.head  # written to durable, external storage
...
trail.verify(expected_head=anchor)  # False if anything was removed

Nor can a chain alone catch a rebuild: rewrite a record and recompute every hash after it, and the chain is consistent again. That is what keys is for. With it, every record must carry a signature that verifies under the public key its key_id names, so a rebuilt chain fails unless it was signed with a key the verifier trusts.

Parameters:

Name Type Description Default
expected_head str | None

The chain head recorded out-of-band. When given, the trail must also end on this hash. Omit it and truncation goes undetected — see above.

None
keys Mapping[str, bytes | str] | None

Trusted public keys as PEM, by key id. When given, an unsigned record, an unknown key id, or a bad signature fails verification.

None

Returns:

Type Description
bool

True if the chain is intact, ends at expected_head when one

bool

was supplied, and every record is validly signed when keys was.

Source code in .sdk/src/tulip/control/audit.py
def verify(
    self,
    *,
    expected_head: str | None = None,
    keys: Mapping[str, bytes | str] | None = None,
) -> bool:
    """Whether the chain is internally consistent, optionally un-truncated and signed.

    On its own this catches every edit, reorder, and deletion **from the
    middle** of the chain: each of those leaves a record whose stored hash
    no longer matches its contents, or whose ``prev_hash`` no longer points
    at the record before it.

    It cannot, on its own, catch a *truncation*. Dropping records from the
    end — or discarding the trail entirely — leaves a shorter chain that is
    perfectly valid on its own terms, so this returns ``True``. That is a
    property of hash chains in general, not of this implementation: nothing
    inside a chain can attest to a link that was never handed to it.

    Truncation is what ``expected_head`` is for. Persist :attr:`head`
    somewhere the agent cannot reach — a WORM bucket, a append-only log, a
    transparency log, a co-signer — and pass it back here. Every attack
    above, truncation included, changes the head:

    ```python
    anchor = trail.head  # written to durable, external storage
    ...
    trail.verify(expected_head=anchor)  # False if anything was removed
    ```

    Nor can a chain alone catch a *rebuild*: rewrite a record and recompute
    every hash after it, and the chain is consistent again. That is what
    ``keys`` is for. With it, every record must carry a signature that
    verifies under the public key its ``key_id`` names, so a rebuilt chain
    fails unless it was signed with a key the verifier trusts.

    Args:
        expected_head: The chain head recorded out-of-band. When given, the
            trail must also *end* on this hash. Omit it and truncation goes
            undetected — see above.
        keys: Trusted public keys as PEM, by key id. When given, an unsigned
            record, an unknown key id, or a bad signature fails verification.

    Returns:
        ``True`` if the chain is intact, ends at ``expected_head`` when one
        was supplied, and every record is validly signed when ``keys`` was.
    """
    prev = _GENESIS
    for i, rec in enumerate(self._records):
        if rec.seq != i or rec.prev_hash != prev:
            return False
        if _entry_hash(rec.seq, rec.ts, rec.event_type, rec.payload, rec.prev_hash) != rec.hash:
            return False
        prev = rec.hash
    if expected_head is not None and self.head != expected_head:
        return False
    if keys is not None:
        return _signatures_valid(self._records, keys)
    return True

export_jsonl

export_jsonl() -> str

The chain as newline-delimited JSON — one record per line, SIEM-ready.

Source code in .sdk/src/tulip/control/audit.py
def export_jsonl(self) -> str:
    """The chain as newline-delimited JSON — one record per line, SIEM-ready."""
    return "\n".join(json.dumps(_exported(rec), default=str) for rec in self._records)

from_records classmethod

from_records(records: Iterable[AuditRecord]) -> AuditTrail

Rebuild a trail from records (e.g. to :meth:verify an exported chain).

Source code in .sdk/src/tulip/control/audit.py
@classmethod
def from_records(cls, records: Iterable[AuditRecord]) -> AuditTrail:
    """Rebuild a trail from records (e.g. to :meth:`verify` an exported chain)."""
    trail = cls()
    trail._records = list(records)
    return trail

AuditRecord dataclass

AuditRecord(seq: int, ts: str, event_type: str, payload: dict[str, Any], prev_hash: str, hash: str, key_id: str | None = None, signature: str | None = None)

One link in the audit chain. hash commits to prev_hash.

AuditSigner

Bases: Protocol

Signs a record's hash. key_id names the key a verifier should use.

sign

sign(data: bytes) -> bytes

A signature over data.

Source code in .sdk/src/tulip/control/audit.py
def sign(self, data: bytes) -> bytes:
    """A signature over ``data``."""
    ...

Ed25519Signer

Ed25519Signer(private_key: Any, *, key_id: str | None = None)

An Ed25519 :class:AuditSigner.

Parameters:

Name Type Description Default
private_key Any

A cryptography Ed25519 private key.

required
key_id str | None

The name verifiers look the public key up by. Defaults to the first 16 hex characters of the SHA-256 of the raw public key, so the same key always gets the same id.

None
Source code in .sdk/src/tulip/control/audit.py
def __init__(self, private_key: Any, *, key_id: str | None = None) -> None:
    serialization, _, _ = _crypto()
    self._key = private_key
    raw = private_key.public_key().public_bytes(
        serialization.Encoding.Raw, serialization.PublicFormat.Raw
    )
    self.key_id: str = key_id or hashlib.sha256(raw).hexdigest()[:16]

generate classmethod

generate(*, key_id: str | None = None) -> Ed25519Signer

A signer with a new random key.

Source code in .sdk/src/tulip/control/audit.py
@classmethod
def generate(cls, *, key_id: str | None = None) -> Ed25519Signer:
    """A signer with a new random key."""
    _, ed25519, _ = _crypto()
    return cls(ed25519.Ed25519PrivateKey.generate(), key_id=key_id)

from_pem classmethod

from_pem(pem: bytes | str, *, password: bytes | None = None, key_id: str | None = None) -> Ed25519Signer

A signer from a PEM-encoded (PKCS#8) Ed25519 private key.

Source code in .sdk/src/tulip/control/audit.py
@classmethod
def from_pem(
    cls, pem: bytes | str, *, password: bytes | None = None, key_id: str | None = None
) -> Ed25519Signer:
    """A signer from a PEM-encoded (PKCS#8) Ed25519 private key."""
    serialization, ed25519, _ = _crypto()
    data = pem.encode("ascii") if isinstance(pem, str) else pem
    key = serialization.load_pem_private_key(data, password=password)
    if not isinstance(key, ed25519.Ed25519PrivateKey):
        raise TypeError("expected an Ed25519 private key")
    return cls(key, key_id=key_id)

private_key_pem

private_key_pem(*, password: bytes | None = None) -> bytes

The private key as PKCS#8 PEM, encrypted when password is given.

Source code in .sdk/src/tulip/control/audit.py
def private_key_pem(self, *, password: bytes | None = None) -> bytes:
    """The private key as PKCS#8 PEM, encrypted when ``password`` is given."""
    serialization, _, _ = _crypto()
    encryption = (
        serialization.BestAvailableEncryption(password)
        if password
        else serialization.NoEncryption()
    )
    pem: bytes = self._key.private_bytes(
        serialization.Encoding.PEM, serialization.PrivateFormat.PKCS8, encryption
    )
    return pem

public_key_pem

public_key_pem() -> bytes

The public key as SubjectPublicKeyInfo PEM — what a verifier needs.

Source code in .sdk/src/tulip/control/audit.py
def public_key_pem(self) -> bytes:
    """The public key as SubjectPublicKeyInfo PEM — what a verifier needs."""
    serialization, _, _ = _crypto()
    pem: bytes = self._key.public_key().public_bytes(
        serialization.Encoding.PEM, serialization.PublicFormat.SubjectPublicKeyInfo
    )
    return pem

verify_jsonl

verify_jsonl(text: str, *, keys: Mapping[str, bytes | str] | None = None, expected_head: str | None = None) -> bool

Verify an exported trail from its JSONL alone.

Needs no Tulip runtime state: an auditor with the export and the public keys runs this and nothing else. A line that is not a record fails.

Source code in .sdk/src/tulip/control/audit.py
def verify_jsonl(
    text: str,
    *,
    keys: Mapping[str, bytes | str] | None = None,
    expected_head: str | None = None,
) -> bool:
    """Verify an exported trail from its JSONL alone.

    Needs no Tulip runtime state: an auditor with the export and the public
    keys runs this and nothing else. A line that is not a record fails.
    """
    records: list[AuditRecord] = []
    for line in text.splitlines():
        if not line.strip():
            continue
        try:
            records.append(AuditRecord(**json.loads(line)))
        except (TypeError, ValueError):
            return False
    return AuditTrail.from_records(records).verify(expected_head=expected_head, keys=keys)

AuditHook

AuditHook(trail: AuditTrail, *, priority: int = HookPriority.OBSERVABILITY_DEFAULT)

Bases: HookProvider

Records the agent's lifecycle into a tamper-evident :class:AuditTrail.

Source code in .sdk/src/tulip/control/governed.py
def __init__(
    self,
    trail: AuditTrail,
    *,
    priority: int = HookPriority.OBSERVABILITY_DEFAULT,
) -> None:
    self._trail = trail
    self._priority = priority

name property

name: str

Hook provider name for identification.

on_iteration_start async

on_iteration_start(iteration: int, state: AgentState) -> None

Called at the start of each agent iteration.

Parameters:

Name Type Description Default
iteration int

Current iteration number (0-indexed)

required
state AgentState

Current agent state

required
Source code in .sdk/src/tulip/hooks/provider.py
async def on_iteration_start(
    self,
    iteration: int,
    state: AgentState,
) -> None:
    """Called at the start of each agent iteration.

    Args:
        iteration: Current iteration number (0-indexed)
        state: Current agent state
    """

on_iteration_end async

on_iteration_end(iteration: int, state: AgentState) -> None

Called at the end of each agent iteration.

Parameters:

Name Type Description Default
iteration int

Current iteration number (0-indexed)

required
state AgentState

Current agent state

required
Source code in .sdk/src/tulip/hooks/provider.py
async def on_iteration_end(
    self,
    iteration: int,
    state: AgentState,
) -> None:
    """Called at the end of each agent iteration.

    Args:
        iteration: Current iteration number (0-indexed)
        state: Current agent state
    """

on_before_model_call async

on_before_model_call(event: BeforeModelCallEvent) -> None

Called before each model.complete() call.

Modify event.messages to change what the model sees. event.tools is read-only (inspect only).

Parameters:

Name Type Description Default
event BeforeModelCallEvent

Write-protected event. Writable: messages.

required
Source code in .sdk/src/tulip/hooks/provider.py
async def on_before_model_call(
    self,
    event: BeforeModelCallEvent,
) -> None:
    """Called before each model.complete() call.

    Modify event.messages to change what the model sees.
    event.tools is read-only (inspect only).

    Args:
        event: Write-protected event. Writable: messages.
    """

on_after_model_call async

on_after_model_call(event: AfterModelCallEvent) -> None

Called after each model.complete() call.

Set event.retry = True to discard response and re-call. Set event.response to replace the response. event.messages is read-only.

Parameters:

Name Type Description Default
event AfterModelCallEvent

Write-protected event. Writable: response, retry.

required
Source code in .sdk/src/tulip/hooks/provider.py
async def on_after_model_call(
    self,
    event: AfterModelCallEvent,
) -> None:
    """Called after each model.complete() call.

    Set event.retry = True to discard response and re-call.
    Set event.response to replace the response.
    event.messages is read-only.

    Args:
        event: Write-protected event. Writable: response, retry.
    """

Governed agents

An Agent pre-wired with grounding, guardrails, and an audit trail.

governed_agent

governed_agent(model: Any = None, tools: list[Any] | None = None, *, system_prompt: str | None = None, profile: GovernanceProfile | None = None, audit_trail: AuditTrail | None = None, hooks: list[Any] | None = None, **kwargs: Any) -> GovernedAgent

Build a secure-by-default agent: grounded, guarded, and audited.

Parameters:

Name Type Description Default
model Any

Model string or instance (as :class:tulip.Agent).

None
tools list[Any] | None

Tools available to the agent.

None
system_prompt str | None

System prompt.

None
profile GovernanceProfile | None

Which controls to enable (default: all on).

None
audit_trail AuditTrail | None

Reuse an existing trail; one is created if omitted.

None
hooks list[Any] | None

Extra hooks to add alongside the security hooks.

None
**kwargs Any

Passed through to :class:tulip.Agent.

{}

Returns:

Name Type Description
A GovernedAgent

class:GovernedAgent wrapping the configured agent and its audit trail.

Source code in .sdk/src/tulip/control/governed.py
def governed_agent(
    model: Any = None,
    tools: list[Any] | None = None,
    *,
    system_prompt: str | None = None,
    profile: GovernanceProfile | None = None,
    audit_trail: AuditTrail | None = None,
    hooks: list[Any] | None = None,
    **kwargs: Any,
) -> GovernedAgent:
    """Build a secure-by-default agent: grounded, guarded, and audited.

    Args:
        model: Model string or instance (as :class:`tulip.Agent`).
        tools: Tools available to the agent.
        system_prompt: System prompt.
        profile: Which controls to enable (default: all on).
        audit_trail: Reuse an existing trail; one is created if omitted.
        hooks: Extra hooks to add alongside the security hooks.
        **kwargs: Passed through to :class:`tulip.Agent`.

    Returns:
        A :class:`GovernedAgent` wrapping the configured agent and its audit trail.
    """
    profile = profile or GovernanceProfile()
    # NB: an empty AuditTrail is falsy (len 0), so check identity, not truthiness.
    trail = audit_trail if audit_trail is not None else AuditTrail()
    hook_list: list[Any] = list(hooks or [])
    if profile.guardrails:
        hook_list.append(GuardrailsHook())
    if profile.audit:
        hook_list.append(AuditHook(trail))
    agent = Agent(
        model=model,
        tools=tools,
        system_prompt=system_prompt,
        grounding=profile.grounding,
        hooks=hook_list,
        **kwargs,
    )
    return GovernedAgent(agent=agent, audit_trail=trail, profile=profile)

GovernedAgent dataclass

GovernedAgent(agent: Agent, audit_trail: AuditTrail, profile: GovernanceProfile)

A secure-by-default :class:tulip.Agent plus its audit trail.

run / run_sync pass through to the wrapped agent; audit_trail is the tamper-evident record of everything it did.

arun async

arun(prompt: str, **kwargs: Any) -> Any

Async, thread-free twin of run_sync — delegates to the wrapped agent's arun so a governed agent runs where threads aren't available (e.g. the browser / Pyodide).

Source code in .sdk/src/tulip/control/governed.py
async def arun(self, prompt: str, **kwargs: Any) -> Any:
    """Async, thread-free twin of ``run_sync`` — delegates to the wrapped
    agent's ``arun`` so a governed agent runs where threads aren't
    available (e.g. the browser / Pyodide)."""
    return await self.agent.arun(prompt, **kwargs)

GovernanceProfile dataclass

GovernanceProfile(grounding: bool = True, guardrails: bool = True, audit: bool = True)

Which secure-by-default controls a :func:governed_agent turns on.

All on by default — that is what makes the agent secure out of the box.

Verification

Evidence quality and adversarial refutation, feeding the require_verification_score and min_severity rules on a policy.

verify async

verify(finding: FindingLike, *, skeptics: Sequence[Skeptic] | None = None, threshold: float = 0.6) -> VerificationResult

Independently challenge a finding; return whether it survives.

Runs each skeptic (default: a single :class:EvidenceQualitySkeptic), collects their refutations, and re-grades confidence as the grounding score minus the refutation penalties — where non-fatal penalties are capped (:data:_MAX_NONFATAL_PENALTY) so volume of caveats alone can't refute a well-grounded finding; a single fatal refutation zeroes it outright. A finding survives only if nothing fatal was raised and confidence clears threshold.

Parameters:

Name Type Description Default
finding FindingLike

A :class:~tulip.control.findings.Evidence or finding-shaped mapping (framework-agnostic).

required
skeptics Sequence[Skeptic] | None

The challenge panel; defaults to the deterministic skeptic. Plug semantic/LLM skeptics here.

None
threshold float

Minimum confidence to survive (default 0.6).

0.6

Returns:

Name Type Description
A VerificationResult

class:VerificationResult.

Source code in .sdk/src/tulip/control/verification.py
async def verify(
    finding: FindingLike,
    *,
    skeptics: Sequence[Skeptic] | None = None,
    threshold: float = 0.6,
) -> VerificationResult:
    """Independently challenge a finding; return whether it survives.

    Runs each skeptic (default: a single :class:`EvidenceQualitySkeptic`),
    collects their refutations, and re-grades confidence as the grounding score
    minus the refutation penalties — where non-fatal penalties are **capped**
    (:data:`_MAX_NONFATAL_PENALTY`) so volume of caveats alone can't refute a
    well-grounded finding; a single ``fatal`` refutation zeroes it outright. A
    finding survives only if nothing fatal was raised and confidence clears
    ``threshold``.

    Args:
        finding: A :class:`~tulip.control.findings.Evidence` or finding-shaped
            mapping (framework-agnostic).
        skeptics: The challenge panel; defaults to the deterministic skeptic.
            Plug semantic/LLM skeptics here.
        threshold: Minimum confidence to survive (default 0.6).

    Returns:
        A :class:`VerificationResult`.
    """
    panel: list[Skeptic] = list(skeptics) if skeptics is not None else [EvidenceQualitySkeptic()]
    refutations: list[Refutation] = []
    for skeptic in panel:
        refutations.extend(await skeptic.challenge(finding))

    base = _coerce(finding).gsar_score
    fatal = any(r.weight == "fatal" for r in refutations)
    nonfatal = sum(_PENALTY.get(r.weight, 0.2) for r in refutations if r.weight != "fatal")
    confidence = 0.0 if fatal else max(0.0, min(1.0, base - min(nonfatal, _MAX_NONFATAL_PENALTY)))

    survives = not fatal and confidence >= threshold
    notes = (
        "Survives independent challenge."
        if survives
        else "Refuted by independent challenge — do not act on this finding as-is."
    )
    return VerificationResult(
        survives=survives,
        confidence=confidence,
        evidence_quality=confidence,
        refutations=refutations,
        alternatives=[],
        notes=notes,
    )

VerificationResult dataclass

VerificationResult(survives: bool, confidence: float, evidence_quality: float, refutations: list[Refutation] = list(), alternatives: list[str] = list(), notes: str = '')

The outcome of verifying a finding.

survives is False if any refutation is fatal or confidence falls below the threshold. alternatives is populated by semantic skeptics (the deterministic one leaves it empty).

Evidence

Bases: BaseModel

A grounded security finding.

The gsar_score and evidence_refs fields are required: a Evidence always knows how strongly it is grounded and what it is grounded in. Build findings via :func:tulip.control.ground_finding rather than constructing them directly — that is the path that enforces the grounding threshold.

Severity

Bases: StrEnum

Ordered severity band. StrEnum so it serialises as the bare string.

Not directly comparable with < (string ordering would be wrong); use :func:severity_at_least or :data:SEVERITY_ORDER for ranking.