Skip to Content

AI Gateway API

The Auris AI Gateway exposes a set of API endpoints that power the built-in administrative assistant available in the Console. The assistant can answer questions about your tenant configuration, produce structured action plans, execute read operations autonomously, and request human approval before taking any write action.

All AI traffic is routed through a LiteLLM proxy container (litellm:4000) that normalises provider APIs, enforces per-application virtual key budgets, and provides a unified model identifier (altovar-default). Swapping the underlying model or provider is a configuration change — no application code changes are required.

All endpoints in this group require the admin:all permission and the x-tenant header.

The AI Gateway can be disabled at the infrastructure level by setting the environment variable AURIS_AI_ENABLED=false. When disabled, the /api/ai/chat and /api/ai/explain endpoints return HTTP 503. Note: the approval routes (GET /api/ai/approvals/[id] and POST /api/ai/approvals/[id]/decision) do not check AURIS_AI_ENABLED and will continue to respond normally even when AI is disabled.


Chat

POST/api/ai/chatRequires: admin:all

Send a conversational message to the AI assistant. The assistant processes the full message history, optionally calls read-only tools to look up live tenant data, and returns a reply. Write actions are never executed directly — the assistant raises an approval request that must be confirmed by the caller before the action runs.

Request body

FieldTypeRequiredDescription
messagesMessage[]YesConversation history. Min 1, max 40 items. Each content max 4 000 chars.
modestringNoAssistant mode: info, plan, agent, or bypass. Default: info.
effortstringNoResponse quality/cost level: low, medium, high, or max. Default: medium.
reasoningbooleanNoOverride the effort-derived thinking toggle.
contextstringNoCurrent page label (max 500 chars) injected into the system prompt.
localestringNoBCP 47 locale tag used to localise the assistant’s reply.
customInstructionsstringNoAdditional instructions appended to the system prompt (max 4 000 chars).
platformPolicyobjectNoRestrict the assistant’s capabilities for this call. See Platform Policy.
approvalobjectNoApproval decision for a pending action. See Approval Workflow.
contextTagsContextTag[]NoUp to 20 structured context tags ({id, label, type?, value?}).
skillIdstringNoActivate a named skill loaded by the assistant runtime.
attachmentsAttachment[]NoUp to 10 file attachments. Each max 25 MB ({name, size, type, dataBase64}).
questionAnswerobjectNoReply to a clarifying question raised in the previous turn ({questionId, answer}).

Message object

{ "role": "user", "content": "How many active sessions does [email protected] have?" }
FieldTypeValues
rolestringuser | assistant
contentstringMessage text, max 4 000 chars

Response body

{ "reply": "Alice currently has 3 active sessions across 2 devices.", "reasoning": "...", "reasoningDurationMs": 1240, "toolCalls": [ { "id": "call_abc", "name": "list_sessions", "arguments": "{\"userId\":\"usr_abc123\"}", "result": "{\"sessions\":[\"...\"]}", "durationMs": 182 } ] }
FieldTypeDescription
replystringThe assistant’s text response.
reasoningstringInternal chain-of-thought, present only when thinking is enabled (effort high/max or reasoning: true).
reasoningDurationMsnumberTime spent in the thinking phase (ms).
actionRequestobjectPresent when the assistant has identified a write action requiring approval. See Action Request.
actionResultobjectPresent when an approved action has been executed. Fields: actionId, actionName, status (executed|rejected|failed), details (string, optional).
stepUpobjectPresent (HTTP 428) when the action requires step-up authentication before it can proceed. See Step-Up Authentication.
questionobjectA clarifying question from the model. Fields: id, text, suggestedAnswers?, required?. Resend with questionAnswer to answer.
toolCallsobject[]Read-tool invocations made during this turn. arguments and result are JSON-serialised strings, not parsed objects.

Modes

The mode field changes the assistant’s behaviour and prompt strategy.

ModeBehaviour
infoDefault. Answers questions, explains concepts, and looks up live data. Does not propose actions.
planProduces a numbered execution plan with prerequisites, risk notes, and a rollback strategy. Useful before making changes.
agentAction-oriented. Distinguishes between read-safe operations the assistant can perform immediately and write operations that require explicit approval.
bypassElevated operator mode for full-platform orchestration. Blocked unless the environment variable AURIS_AI_ALLOW_BYPASS=true is set. Returns HTTP 403 otherwise.

Effort levels

The effort field controls the model’s token budget, temperature, and whether chain-of-thought reasoning is enabled.

LevelMax tokensTemperatureReasoning effortThinking enabled
low1 2000.6minimalNo
medium2 2000.7minimalNo
high3 2000.75mediumYes
max4 0960.8highYes

Setting reasoning: false forces reasoning effort to minimal regardless of effort. Setting reasoning: true raises it to at least medium.

Tool loop

In agent mode the assistant has access to a set of read-only tools it can invoke autonomously (up to 5 rounds per request). Write actions are never invoked by the model directly — they are always surfaced as an approval request.

Read tools (no approval required):

search_users, get_user, get_user_roles, get_user_groups, list_audit_events, count_active_sessions, list_sessions, list_roles, get_role, list_groups, list_applications, get_application, check_url_connectivity, list_ip_rules, list_webhooks, get_security_defaults, get_tenant_info, list_security_alerts, get_mfa_stats, list_trusted_devices, list_organizations, get_organization, list_sso_connections, list_email_templates, fga_check, fga_list_tuples, list_log_streams, get_log_stream, list_pending_events

Write actions (always require approval):

user_disable, user_enable, session_revoke, session_revoke_all, user_force_password_reset, mfa_reset, ip_rule_block, ip_rule_disable, ip_rule_enable, ip_rule_delete, create_ip_allow_rule, send_password_reset_email, revoke_trusted_device, revoke_all_trusted_devices, acknowledge_security_alert, resolve_security_alert, test_log_stream, enable_log_stream, disable_log_stream, delete_log_stream, application_disable, application_enable, assign_role_to_user, remove_role_from_user, create_user, create_application, create_webhook, enable_webhook, disable_webhook, create_organization, update_organization, create_role, create_group, add_user_to_group, remove_user_from_group, delete_group, add_organization_member, remove_organization_member, update_organization_member_role, send_verification_email, update_security_defaults, update_user, update_application, regenerate_client_secret, delete_user, delete_application, delete_webhook, delete_organization, delete_role, fga_write_tuple, fga_delete_tuple

Platform policy

The optional platformPolicy field scopes down what the assistant is allowed to do for this call. This is useful when embedding the assistant in a context where only a subset of capabilities should be available.

{ "platformPolicy": { "name": "read-only-audit", "capabilities": ["list_audit_events", "get_user"], "allowRead": true, "allowWrite": false, "permissionMode": "strict" } }
FieldTypeDescription
namestringLabel for the policy (used in audit logs).
capabilitiesstring[]Allowlist of tool names available to the model. Omit to allow all.
allowReadbooleanWhether read tools may be called.
allowWritebooleanWhether write-action approval requests may be raised.
permissionModestringask — prompt user for confirmation; strict — deny any unlisted capability; none — no enforcement.

Clarifying questions

The model may embed a fenced block in its reply to ask a clarifying question before proceeding. The server strips the block from the visible reply and surfaces it in the question field.

The model appends to its text reply: ```clarify {"id":"q-env","text":"Which environment should I target?","suggestedAnswers":["staging","production"]}
The client should display the question to the operator and re-send the conversation with `questionAnswer: { questionId: "q-env", answer: "production" }`. --- ## Explain <ApiEndpoint method="POST" path="/api/ai/explain" permission="admin:all"> Ask the assistant to explain an opaque object in plain language — for example a raw audit log entry, a token exchange request, a CIBA callback, or a security event. The object is automatically redacted of sensitive keys before being sent to the model. </ApiEndpoint> ### Request body ```json { "topic": "token-exchange", "payload": { "grant_type": "urn:ietf:params:oauth:grant-type:token-exchange", "subject_token_type": "urn:ietf:params:oauth:token-type:access_token", "requested_token_type": "urn:ietf:params:oauth:token-type:refresh_token" }, "context": { "clientId": "app_billing" }, "locale": "en" }
FieldTypeRequiredDescription
topicstringNoHint for the model: risk-assessment, ciba-request, token-exchange, device-flow, or general. Default: general.
payloadobjectYesThe raw object to explain.
contextobjectNoOptional surrounding context (e.g., client metadata).
localestringNoBCP 47 locale.

The following keys are automatically redacted from payload before being sent to the model: token, accessToken, idToken, refreshToken, authorization, password, secret, clientSecret. Redacted values are replaced with the string [REDACTED].

Response body

{ "success": true, "data": { "summary": "This is a standard RFC 8693 token exchange request. The caller is trading an existing access token for a refresh token.", "findings": [ "The subject token type is an access token, which is the standard input for this grant.", "The requested token type is a refresh token — this extends the session lifetime." ], "recommendations": [ "Verify that the client is authorised to perform token exchange (check the allowed_grant_types list for app_billing).", "Confirm that the subject token has not expired before processing." ], "confidence": "high" } }
FieldTypeDescription
summarystringOne-paragraph plain-language explanation.
findingsstring[]Key observations extracted from the payload.
recommendationsstring[]Actionable next steps.
confidencestringModel confidence in the explanation: low, medium, or high.

Action Approvals

When the assistant identifies a write action in the conversation, it does not execute the action immediately. Instead it creates an AiActionApproval record and returns an actionRequest object in the chat response. The operator must explicitly approve or reject the action before it runs.

This design means the assistant can never modify tenant state without a human decision in the loop.

Approval lifecycle

pending --> approved --> executed --> failed --> rejected --> expired (after 10 minutes)

Approval records are retained for 30 days after expiry, then deleted.

Action risk levels

Every write action is assigned a risk level that determines whether step-up authentication is required.

RiskStep-up requiredExample actions
lowNosend_password_reset_email, acknowledge_security_alert, test_log_stream
mediumNouser_enable, session_revoke, create_user, add_user_to_group
highNouser_disable, session_revoke_all, delete_user, delete_organization, fga_write_tuple
criticalYes (AAL2)Actions classified critical require requiredAcr: aal2 step-up before execution.

Action request object

When a write action is identified, the chat response includes an actionRequest object:

{ "actionRequest": { "actionId": "apr_hk3mn9xq", "actionName": "user_disable", "summary": "Disable the account for [email protected] (usr_abc123).", "risk": "high", "requiresApproval": true, "requiresStepUp": false, "requiredAcr": null, "expiresAt": "2025-02-18T10:10:00Z", "approvalToken": "eyJ..." } }

To approve the action inline, resend the conversation with the approval field:

{ "messages": [...], "approval": { "actionId": "apr_hk3mn9xq", "approvalToken": "eyJ...", "approved": true, "decisionReason": "Confirmed with the account owner." } }

approvalToken is optional in the inline approval object.

Alternatively, use the dedicated approval endpoints below.


Get Approval

GET/api/ai/approvals/[id]Requires: admin:all

Retrieve the current state of a pending or completed approval record.

Success response

{ "actionId": "apr_hk3mn9xq", "actionName": "user_disable", "summary": "Disable the account for [email protected].", "risk": "high", "status": "pending", "requiresStepUp": false, "requiredAcr": null, "expiresAt": "2025-02-18T10:10:00Z", "createdAt": "2025-02-18T10:00:00Z", "approvedAt": null, "rejectedAt": null, "executedAt": null, "decisionReason": null, "actionResultStatus": null, "actionResultDetails": null }

The response is the ActionApprovalView object returned directly — there is no ok/data envelope.

Error response (404)

{ "error": "Approval request not found" }

Error codes

HTTPDescription
404No approval record with this ID.

Submit Decision

POST/api/ai/approvals/[id]/decisionRequires: admin:all

Approve or reject a pending action. If the action requires step-up authentication, the server returns HTTP 428 with a challenge. Complete the challenge and re-submit with the step-up response to proceed.

Request body

{ "approved": true, "decisionReason": "Verified with the account owner before proceeding.", "stepUp": { "challengeId": "chal_xyz789", "method": "totp", "response": "482910" } }
FieldTypeRequiredDescription
approvedbooleanYestrue to approve, false to reject.
decisionReasonstringNoFreeform reason logged to the audit trail (max 2 000 chars).
stepUpobjectConditionalStep-up response. Required if the server previously returned HTTP 428 for this approval.

Success response

{ "status": "executed", "approval": { "actionId": "apr_hk3mn9xq", "actionName": "user_disable", "...": "ActionApprovalView fields" }, "actionResult": { "actionId": "apr_hk3mn9xq", "actionName": "user_disable", "status": "executed", "details": "User usr_abc123 disabled successfully." } }

The response is the ApprovalDecisionResult object returned directly — there is no ok/data envelope. actionResult.details is a plain string (or absent), not a structured object.

The status field in the response reflects the outcome:

StatusMeaning
approvedDecision recorded; action is queued for execution (async).
rejectedDecision recorded; action will not run.
executedAction ran successfully in this call.
failedAction was approved but execution failed. Check actionResult.details.
step_up_requiredHTTP 428. Step-up challenge issued. Re-submit with stepUp fields.

Error codes

HTTPDescriptionResponse body
404Approval record does not exist.{ "error": "..." }
410The 10-minute TTL has elapsed. The action cannot be executed.{ "error": "..." }
403Step-up challenge response was invalid or expired.{ "error": "<message>" }

Step-Up Authentication

Critical-risk actions require the operator to complete an MFA challenge before the action executes. This is separate from the session-level MFA and provides a fresh proof of presence.

Flow

  1. The operator calls POST /api/ai/approvals/[id]/decision with approved: true but without stepUp fields.
  2. If the action requires step-up (requiresStepUp: true), the server returns HTTP 428:
{ "status": "step_up_required", "approval": { "actionId": "apr_hk3mn9xq", "...": "ActionApprovalView fields" }, "stepUp": { "challengeId": "chal_xyz789", "availableMethods": ["totp", "webauthn"], "expiresIn": 300, "requiredAcr": "aal2" } }

There is no ok/error envelope — the body is the ApprovalDecisionResult object with status: "step_up_required" directly.

  1. The client presents the MFA challenge to the operator.
  2. The operator completes the challenge. The client re-submits the decision with the stepUp object:
{ "approved": true, "stepUp": { "challengeId": "chal_xyz789", "method": "totp", "response": "482910" } }
  1. The server validates the response and, if successful, executes the action.

ACR level hierarchy

LevelNumeric valueDescription
aal11Password only
aal22Password + second factor (TOTP, WebAuthn, SMS)
aal33Password + hardware-bound second factor

The operator’s session ACR must meet or exceed requiredAcr after step-up.


Audit Events

Every state transition in the approval lifecycle emits an audit event visible in the Audit Logs API.

EventTrigger
AI_ACTION_APPROVAL_REQUESTEDAn approval record was created.
AI_ACTION_STEP_UP_REQUIREDHTTP 428 was returned; challenge issued.
AI_ACTION_STEP_UP_VERIFIEDStep-up challenge passed.
AI_ACTION_APPROVEDOperator approved the action.
AI_ACTION_REJECTEDOperator rejected the action.
AI_ACTION_EXECUTEDAction ran successfully.
AI_ACTION_FAILEDAction was approved but execution failed.
AI_ACTION_APPROVAL_EXPIREDApproval record reached its 10-minute TTL without a decision.

Permissions

All endpoints in this group require the admin:all permission. There is no read-only scoping — the assistant’s ability to query live tenant data and raise approval requests for write actions is considered an administrative capability.

The bypass mode (mode: "bypass") requires the AURIS_AI_ALLOW_BYPASS=true environment variable in addition to the admin:all permission. Requests with mode: "bypass" return HTTP 403 unless this flag is explicitly set.


Environment Variables

VariableDefaultDescription
AURIS_AI_ENABLEDtrueSet to false to disable AI on /api/ai/chat and /api/ai/explain (returns 503). The approvals routes are unaffected.
AURIS_AI_API_KEY—Provider API key. If absent, /api/ai/chat and /api/ai/explain return HTTP 503 with code AI_DISABLED — same effect as AURIS_AI_ENABLED=false. Also consumed by the LiteLLM gateway.
AURIS_AI_MODELMiniMax-M2.7Model identifier forwarded to the LiteLLM gateway.
AURIS_AI_API_URLhttps://api.minimax.io/v1/text/chatcompletion_v2Provider base URL. Set to http://litellm:4000/v1/chat/completions to route through the internal LiteLLM proxy.
AURIS_AI_TIMEOUT_MS30000Per-request timeout in milliseconds. Exceeded requests return HTTP 504.
AURIS_AI_SYSTEM_INSTRUCTIONS—Additional instructions appended to the system prompt of every chat request.
AURIS_AI_ALLOW_BYPASSfalseMust be true to permit mode: "bypass" requests.