Overview
In this guide, you can learn how to invoke a deployed agent on the MongoDB Atlas Agent Engine. The guide shows how to call the invocation API, invoke an agent from the CLI, forward custom headers to your agent, and resume suspended executions.
To generate the API key or service account credentials that these requests use, see Manage API Keys and Service Accounts.
Invocation API
The invocation API is the API that clients use to call a deployed agent. The Atlas Agent Engine exposes these endpoints on behalf of the agent, so your agent code does not need to define any HTTP routes or start a web server.
Invoke Your Agent
To invoke an agent, send a POST request to the /api/v1/projects/{project_id}/workspaces/{workspace_id}/invoke API endpoint. The response returns the execution result and the execution status. The platform returns the session ID in the X-Session-ID response header.
The following curl example uses these placeholders:
$API_KEY: Your Atlas Agent Engine API key$PROJECT_ID: Your project ID$WORKSPACE_ID: Your workspace ID
Select the tab for your preferred language to view a sample invoke request:
curl -s "https://agentengine.mongodb.com/api/v1/projects/$PROJECT_ID/workspaces/$WORKSPACE_ID/invoke" \ -X POST \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"message": "Hello, agent!"}'
import httpx response = httpx.post( f"https://agentengine.mongodb.com/api/v1/projects/{project_id}/workspaces/{workspace_id}/invoke", headers={"Authorization": f"Bearer {api_key}"}, json={"message": "Hello, agent!"}, timeout=60.0, ) result = response.json()
The response resembles the following output:
{ "success": true, "response": "<agent output>", "execution_id": "string", "status": "completed" }
If the execution suspends for human review, the response also includes suspend_reason and suspend_context fields. To learn more, see Suspend, Review, and Resume Lifecycle in the Human-in-the-Loop guide.
Stream Your Agent's Output
To stream an agent's output as it's produced, send a POST request to the /api/v1/projects/{project_id}/workspaces/{workspace_id}/invokeStream API endpoint. The platform returns the response as a stream of Server-Sent Events (SSE) frames. The stream keeps the connection open until the run completes or fails.
Each SSE frame begins with the data: prefix and contains a single JSON object. The following example shows the format of a streaming frame:
data: {"chunk_type": "text", "content": "Hello, agent!", "metadata": {}, "execution_id": "string"}
A frame can include the following fields:
Field | Description |
|---|---|
| Identifies the type of chunk. The following section lists the possible values. |
| The chunk payload. For |
| Structured metadata that describes the chunk. |
| Identifies the execution that produced the chunk. The platform includes this field when the execution ID is available. |
Chunk Types
The following table describes the chunk types that a stream can carry:
Chunk type | Description |
|---|---|
| An increment of the agent's generated output. |
| A progress update that a tool emits during execution. |
| Marks the start of a subagent execution. The platform emits this chunk only when guardrails are disabled. |
| Marks the end of a subagent execution. The platform emits this chunk only when guardrails are disabled. |
| Marks the end of the stream after the run completes. |
| Marks the end of the stream after the run fails. The frame also carries the error message. |
If the run fails, the stream sends an error chunk that includes the error message. Then, the stream closes.
Custom Output for Opted-In Agents
If you set the features.use_custom_parser flag to true in your agent.yaml file, the stream carries only the custom events that the agent's output parser or its emit_custom_event() calls produce. The platform forwards each custom event as a single data: frame, so the frame contains exactly the JSON object that the agent emitted. The stream does not include the standard chunks such as text, step, or done frames, and it closes when the run completes.
The platform delivers custom events only on the invokeStream endpoint. The synchronous invoke endpoint returns the agent's final output instead.
When this feature flag is not enabled, the stream uses the chunk format described earlier in this section. The stream does not deliver custom events, and an emit_custom_event() call in agent code raises an error.
Other Endpoints
The Atlas Agent Engine also exposes the following API endpoints for streaming, polling, resuming, and stopping executions:
Method and Path | Purpose |
|---|---|
| Streams the agent's output incrementally as it is produced by using the Server-Sent Events protocol. |
| Polls execution status and result. |
| Resumes a suspended execution. The request body carries the reviewer decision. To learn more, see Use the API in the Human-in-the-Loop guide. |
| Cancels an in-progress execution, stops its server-side work, and releases the session's compute. To learn more, see Cancel a Session. |
| Interrupts in-flight tool or LLM calls without ending the session. To learn more, see Interrupt a Tool or LLM Call. |
Every executions/** endpoint requires the workspace_id query parameter. The gateway uses this value to route the request to the owning workspace's Orchestration Engine.
Note
Cancelling or interrupting an execution that already reached a terminal status is idempotent. The request response reports that the request made no change.
Execution Status Lifecycle
An execution moves through the statuses pending, running, and then completed or error.
If the execution suspends for human review, it moves through additional statuses before it completes. To learn about the suspend and resume lifecycle, see Suspend, Review, and Resume Lifecycle in the Human-in-the-Loop guide.
If you cancel an execution, it reaches the cancelled status. This status is distinct from completed and error, so API and UI clients can distinguish a deliberately stopped execution from one that finished or failed. To learn more, see Cancel a Session.
Pool-Full Errors
The Atlas Agent Engine does not autoscale agent deployments. The scaling.replicas field in the agent.yaml file sets a fixed sandbox count, and each session reserves one agent sandbox and one tool sandbox for its lifetime. As a result, this field sets the number of sessions that a deployment can serve concurrently.
The scaling.replicas field accepts a value from 1 to 512 and defaults to 4 when you omit it. When every sandbox is reserved, a new invoke request fails with a pool full error.
The 512 concurrent sandbox limit applies to the Orchestration Engine, which is scoped to a project and can serve more than one agent. The scaling.replicas values of all agents in a project count toward the same ceiling. To run more sandboxes than one project allows, distribute your agents across multiple projects.
To keep invoke requests from exhausting the pool, reuse session IDs across requests. Requests that share a session ID reuse one reservation, but requests that omit a session ID use a new pair of sandboxes. A session ID must be 1 to 128 characters long and can contain letters, numbers, underscores (_), and hyphens (-). Pass the session ID in the --session option of the agentengine invoke command, or in the X-Session-ID header of an API request.
If your invoke requests require per-session isolation, you can't reuse a session ID. To serve more sessions at the same time, increase the scaling.replicas value, or reduce the scaling.agent_idle_ttl_seconds and scaling.tool_idle_ttl_seconds values so that idle sessions release their sandboxes sooner.
Because the Atlas Agent Engine snapshots scaling values at build time, you must build and deploy the agent again for a change to take effect. To learn more about these fields, see Agent Contract Reference. To review all limitations that apply during Public Preview, see MongoDB Atlas Agent Engine Limitations.
Forward Custom Headers
In this section, you can learn how to pass custom HTTP headers, such as user IDs, from your application to an agent running on the Atlas Agent Engine. Your agent can then read those headers at runtime by using the get_current_custom_headers() method.
How It Works
On an invoke or resume request, the API Gateway processes HTTP headers with the prefix X-Mdb-Agent-Engine-Custom- by performing the following steps:
The API Gateway extracts the header.
The gateway strips the prefix and makes the header name lowercase. For example,
X-Mdb-Agent-Engine-Custom-Authorizationbecomesauthorization.The gateway forwards the headers to the agent as a dictionary.
Headers are forwarded in-memory through the execution pipeline, never persisted to the database, and discarded by the pipeline when the execution completes.
Note
The Atlas Agent Engine does not persist custom headers. Your application must resend them on every request, including resume requests.
Limits
The following table shows the limits for custom headers:
Limit | Value |
|---|---|
Maximum number of custom headers | 50 |
Maximum size per header | 8 KiB |
Send Custom Headers
Add X-Mdb-Agent-Engine-Custom- prefixed headers to your invoke request. The API Gateway strips the prefix before the agent receives them.
The following examples use the same placeholders as the invocation API examples. They also use placeholders for the custom headers that you want to forward.
Select the tab for your preferred language to see an example invoke request with custom headers:
curl -s "https://agentengine.mongodb.com/api/v1/projects/$PROJECT_ID/workspaces/$WORKSPACE_ID/invoke" \ -X POST \ -H "Authorization: Bearer $API_KEY" \ -H "X-Mdb-Agent-Engine-Custom-Authorization: my-user-id" \ -H "X-Mdb-Agent-Engine-Custom-Tenant-Id: acme-corp" \ -H "Content-Type: application/json" \ -d '{"message": "Hello, agent!"}'
import httpx response = httpx.post( f"https://agentengine.mongodb.com/api/v1/projects/{project_id}/workspaces/{workspace_id}/invoke", headers={ "Authorization": f"Bearer {api_key}", "X-Mdb-Agent-Engine-Custom-Authorization": "my-user-id", "X-Mdb-Agent-Engine-Custom-Tenant-Id": "acme-corp", }, json={"message": "Hello, agent!"}, )
The agent receives the following dictionary after you run the preceding code:
{"authorization": "my-user-id", "tenant-id": "acme-corp"}
Read Headers in Your Agent
Use the get_current_custom_headers() method from agent_engine_runner_shared inside any tool to access the headers that the API Gateway forwarded to the agent. The following Python code shows how to use get_current_custom_headers() to access the forwarded headers:
from agent_engine_sdk_langgraph import App from agent_engine_runner_shared import get_current_custom_headers app = App(app_name="My Agent") def call_external_api(query: str) -> str: """Call an external API using the caller's user ID.""" headers = get_current_custom_headers() user_id = headers.get("authorization", "") tenant = headers.get("tenant-id", "") response = httpx.get( "https://api.example.com/data", headers={"Authorization": user_id, "X-Tenant-Id": tenant}, params={"q": query}, ) return response.text
The get_current_custom_headers() method returns a dict[str, str]. If the gateway did not send any custom headers, the method returns an empty dictionary.
Resume Requests
Forwarding custom headers also works with resume requests. Because the Atlas Agent Engine does not persist custom headers, you must resend the same X-Mdb-Agent-Engine-Custom- headers when you resume a suspended execution. For an example resume request that forwards custom headers, see Use the API in the Human-in-the-Loop guide.
Stop a Running Agent
Cancelling a request from your client closes only your side of the connection. The execution keeps running on the server until you stop it through the Atlas Agent Engine.
The Atlas Agent Engine provides the following ways to stop work that is already running:
Cancel the session to end the execution and release its compute. To learn more, see Cancel a Session.
Stop the run from the Playground UI. To learn more, see Stop a Run in the Playground.
Interrupt a single in-progress tool or LLM call and let the agent continue the run. To learn more, see Interrupt a Tool or LLM Call.
Mark the session finished from your agent code when the agent does not need it anymore. To learn more, see Mark a Session Finished from Your Agent.
Cancel a Session
To cancel an in-progress execution, send a POST request to the /api/v1/projects/{project_id}/executions/{execution_id}/cancel endpoint and set the workspace_id query parameter. The request does not take a body.
Select the tab for your preferred language to see an example of cancelling an execution:
curl -X POST "https://agentengine.mongodb.com/api/v1/projects/$PROJECT_ID/executions/$EXECUTION_ID/cancel?workspace_id=$WORKSPACE_ID" \ -H "Authorization: Bearer $API_KEY"
import httpx response = httpx.post( f"https://agentengine.mongodb.com/api/v1/projects/{project_id}/executions/{execution_id}/cancel", params={"workspace_id": workspace_id}, headers={"Authorization": f"Bearer {api_key}"}, )
The response resembles the following output:
{ "execution_id": "string", "cancelled": true, "status": "cancelled" }
The cancelled field reports whether this request set the execution status to cancelled. The field value is true for a cancelled execution and false when the request did not transition the execution because it had already reached a terminal status, including cancelled from an earlier request. When the field value is false, the status field reports the recorded terminal status instead.
When you cancel an execution, the Atlas Agent Engine does the following:
Cancels in-progress tool or LLM calls and tears down the session's agent sandbox and tool sandbox, and releases the capacity slot. A new execution can reuse the slot.
Cancels every other live execution on the same session, including deep-agent and agent-to-agent executions that are children of the target execution.
Ends an open response stream instead of leaving the stream open until it times out.
Bills the runtime that the session accrued until cancellation.
Note
You cannot resume a cancelled execution. To send another message to the agent, invoke the agent again. The request starts a new execution instead of resuming the cancelled one.
Stop a Run in the Playground
You can also stop a run from the Playground UI instead of calling the cancel API endpoint. Stopping a run from the Playground has the same effect as cancelling the session. To learn more, see Cancel a Session.
To stop a run from the Playground, perform the following steps:
Interrupt a Tool or LLM Call
Interrupting a tool or LLM call stops that call only. The session stays active and the agent continues its run from the interrupted result. Interrupt a call when a single call is stuck or unwanted and you don't want to cancel the whole session.
To interrupt a call, send a POST request to the /api/v1/projects/{project_id}/executions/{execution_id}/interrupt API endpoint. The request takes the following parameters:
Parameter | Type | Required | Description |
|---|---|---|---|
| Query parameter | Yes | The workspace that owns the execution. The Atlas Agent Engine uses this value to route the interrupt request to that workspace's Orchestration Engine. If you omit this parameter, the request fails with a |
| Body field | No | The execution step of the single call to interrupt. If you omit this field, the Atlas Agent Engine interrupts every call that is in progress. |
Select the tab for your preferred language to see an example of interrupting a tool or LLM call:
curl -X POST "https://agentengine.mongodb.com/api/v1/projects/$PROJECT_ID/executions/$EXECUTION_ID/interrupt?workspace_id=$WORKSPACE_ID" \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"step_number": 12}'
import httpx response = httpx.post( f"https://agentengine.mongodb.com/api/v1/projects/{project_id}/executions/{execution_id}/interrupt", params={"workspace_id": workspace_id}, headers={"Authorization": f"Bearer {api_key}"}, json={"step_number": 12}, )
The response resembles the following output:
{ "execution_id": "string", "interrupted": true, "interrupted_steps": [ 12 ], "pending": false, "outcome": "aborted" }
The outcome field describes the result of the interrupt request and returns one of the following values:
Value | Description |
|---|---|
| The Atlas Agent Engine stopped an in-progress call. The |
| There was no in-progress call, so the Atlas Agent Engine applies the interrupt request to the execution's next tool or LLM call. The |
| An unexpired prepared interrupt request already existed. The Atlas Agent Engine did not extend its expiration. |
| The execution already reached a terminal status. |
| The |
Interrupting a call is idempotent and does not change the execution status. If you interrupt every call in a step, the agent ends that turn instead of retrying the calls. If you interrupt only some of the calls in a step, the agent continues normally.
Note
The Atlas Agent Engine can interrupt an I/O-bound call, such as an LLM call or a network tool call. However, a tool that runs a CPU-bound local computation might not stop until it finishes.
Mark a Session Finished from Your Agent
A session holds its agent sandbox and tool sandbox until the idle timeout expires. If you mark a session as finished, your agent can release that compute immediately instead of waiting for the timeout.
Select the tab for your preferred language to see an example of marking a session finished from your agent:
status = app.finish_session()
const status = app.finishSession();
The current turn keeps running and returns its result. After the turn completes, the Atlas Agent Engine cancels any live sub-agent executions and releases the session's pods.
The method returns one of the following values:
Python | TypeScript | Description |
|---|---|---|
|
| The Atlas Agent Engine accepted the request. |
|
| The agent already requested that the Atlas Agent Engine finish this session. |
|
| There is no session to finish. The method returns this value when you call it outside of an agent run, such as from a local script or a tool sandbox, or after the turn ends. |
The method does not raise an error or throw an exception when there is no session to finish.
A turn that suspends for human review or fails keeps its resources so that you can resume and diagnose it. That session falls back to the idle timeout. To learn more about idle timeouts, see Pool-Full Errors.
Invoke an Agent from the CLI
The agentengine invoke command invokes a deployed agent from the terminal without writing any HTTP client code. It reads the workspace from the current directory's .agentengine/state.json file by default.
If you run the command without a message in an interactive terminal, it starts a streaming chat session and reuses the returned session ID across turns. In this interactive mode, the CLI automatically presents an inline review prompt when the agent suspends invocation for human-in-the-loop (HITL) review.
Command Syntax
agentengine invoke [message] [flags] agentengine invoke --file <path> [flags]
The following table describes the available flags:
Flag | Description |
|---|---|
| Stream response chunks as they arrive. To learn how the platform formats streamed output, see Stream Your Agent's Output. |
| Conversation session ID to resume or reuse across turns. |
| User ID to pass to the deployed agent. |
| Read the message from a file instead of a positional argument. |
| JSON metadata object forwarded to the agent alongside the message. You must specify a valid JSON object. This flag can be combined with |
| JSON |
| Output raw JSON, including session ID and streamed chunks. |
| (Monorepo only) Invoke a specific workspace by name. |
| Platform workspace ID, which bypasses local workspace resolution. |
| Platform project ID. |
| Organization ID. |
| Platform API base URL. |
| Named local context from |
| Maximum time to wait for each invoke request. The default value is |
Examples
The following command invokes the agent with a single message:
agentengine invoke "What can you do?"
The following command streams the agent's response as it is produced:
agentengine invoke --stream "Draft a release note"
The following command resumes or continues a named session:
agentengine invoke --session my-session "Follow up question"
The following command reads a message from a file and outputs raw JSON:
cat prompt.txt | agentengine invoke --json
Interactive Payload Commands
When running agentengine invoke in interactive mode, you can manage the payload between turns by using one of the following commands:
Input | Behavior |
|---|---|
| Sets the current payload to the given JSON object. The CLI forwards the payload with each subsequent message until you clear it. |
| Displays the current payload as formatted JSON or prints |
| Clears the current payload. |
Interactive HITL Review
When an agent suspends for human-in-the-loop (HITL) review in interactive mode, the CLI prints the suspend context and prompts you for a decision inline. To learn how the CLI presents the review prompt and how to resume the execution, see Use the CLI in the Human-in-the-Loop guide.
Next Steps
After you invoke an agent, you can monitor your agent's performance and activity. To learn more, see the Monitor guide.