Overview
Atlas Agent Engine guardrails allow project administrators to control the content that passes into and out of an agent's model calls. Administrators set guardrails within a project, and the Orchestration Engine (OE) enforces the guardrails by blocking or modifying matching input and output content.
To apply a guardrail to one or more workspaces, specify the workspaces in the guardrail's workspace_ids field. If you don't specify this field, the guardrail applies to every workspace in the project.
Guardrails complement the Policy Engine. Guardrails control the content that passes into and out of a model, while the Policy Engine controls the actions an agent can take, such as which tools and models it can call.
Guardrail Types
Each guardrail has a type, which determines what content it evaluates, and an action, which determines what happens when the guardrail triggers.
Currently, the Atlas Agent Engine supports only the output_validation guardrail type. You can use guardrails of this type to evaluate an agent's output against a set of regular expression patterns. Define these patterns in the guardrail's config.match_patterns field.
Each guardrail specifies one of the following actions:
Action | Description |
|---|---|
| Denies the call. The execution halts and returns an error to the caller. |
| Removes or transforms the matched content and allows the execution to proceed with the modified content. |
| Records that the guardrail triggered and allows the execution to proceed with the original content. |
| Suspends the execution for human review. The execution proceeds only after a reviewer approves it. |
Guardrail Stages
A guardrail's stage filter determines whether it evaluates content going into the model, coming out of the model, or both. A guardrail with no stage filter applies to all stages.
The following table describes the available stages:
Stage | Evaluates What | Triggers When |
|---|---|---|
| The user content reaching the model | Before a model call |
| The model's response | Before the response returns to the user |
Guardrail Fields
The following table describes the fields that define a guardrail. You can set these fields in the platform UI, or in the JSON request body of a create or update API request.
Field | Possible values | Description |
|---|---|---|
| Any string | (Required) Display name for the guardrail. |
| Any string | Description of the guardrail's purpose. |
|
| (Required) Guardrail type. |
|
| (Required) Action to take when the guardrail triggers. |
|
| Whether the Atlas Agent Engine enforces the guardrail. The Atlas Agent Engine enforces guardrails whose status is |
|
| Array of the stages at which the guardrail evaluates content. To evaluate both stages, list both |
| Array of workspace IDs | Array of the workspaces that the guardrail applies to. To apply the guardrail to every workspace in the project, leave this array empty. |
| Array of regex patterns | Array of the patterns that the guardrail matches content against. Each pattern sets a |
|
| Behavior to apply to content that matches a pattern. This field applies only when |
To view an example request body that sets these fields, see the REST API tab in the following section.
Create and Edit Guardrails
You can create and edit guardrails by using the platform UI or the REST API. To view instructions, select the tab for your preferred method:
Navigate to Manage → Policies, and then select the Guardrails tab.
Then, create a guardrail by clicking the Create guardrail button and configuring the fields in the Create guardrail dialog. To update a guardrail, click the guardrail's Edit icon. To delete a guardrail, click its Delete icon.
Note
The customer-facing REST API exposes guardrail operations. Requests require bearer authentication and appropriate project permissions. To learn more about project roles, see Manage Organizations, Projects, and Workspaces.
The following table lists the available endpoints:
Method | Endpoint | Reference |
|---|---|---|
|
| |
|
| |
|
| |
|
| |
|
|
When calling the create and update endpoints, pass the guardrail's fields as a JSON request body. The following example request body creates a guardrail that redacts any three consecutive digits from the content that reaches the model in a single workspace:
{ "name": "Test guardrail", "description": "Redacts three-digit sequences", "type": "output_validation", "action": "modify", "status": "active", "stage_filter": ["llm_input"], "workspace_ids": ["ws-6a885deac97e9d280fa88cf3"], "config": { "match_patterns": [ { "type": "regex", "value": "\\d{3}" } ], "on_fail": "fix" } }
Guardrail Propagation
After you create, update, or delete a guardrail, the change takes effect on the next execution that the OE handles. A guardrail write reloads the OE's cached guardrails, so the change is effective on the next request.
Guardrail changes do not affect in-flight executions.
Enforcement
During each execution, the OE evaluates the content at each stage against the guardrails that apply to the current project. The following table describes the stages and the corresponding evaluated content:
Stage | Content evaluated |
|---|---|
| The user content is evaluated before a model call. |
| The model's response is evaluated before it returns to the user. |
When a guardrail with a block action triggers, the OE halts the execution and returns an error to the caller. The error identifies the guardrail that caused the block. When a guardrail with a modify action triggers, the execution proceeds with the modified content.
When more than one guardrail matches the same content, the most restrictive action applies. The actions rank in the following order, from most to least restrictive:
blockrequire_reviewmodifylog_only
For example, if one guardrail returns modify and another returns block, the OE blocks the call.
Limitations
Guardrails have the following limitations:
Future executions only: Guardrail changes apply only to executions that start after the change takes effect. Executions already in progress continue with the guardrails they started with.
Content controls only: Guardrails control content that passes into and out of a model. To control the actions an agent can take, such as which tools and models it can call, use the Policy Engine.