AI Gateway API
The Auris AI Gateway exposes a set of API endpoints that power the built-in administrative assistant available in the Console. The assistant can answer questions about your tenant configuration, produce structured action plans, execute read operations autonomously, and request human approval before taking any write action.
All AI traffic is routed through a LiteLLM proxy container (litellm:4000) that normalises provider APIs, enforces per-application virtual key budgets, and provides a unified model identifier (altovar-default). Swapping the underlying model or provider is a configuration change — no application code changes are required.
All endpoints in this group require the admin:all permission and the x-tenant header.
The AI Gateway can be disabled at the infrastructure level by setting the environment variable AURIS_AI_ENABLED=false. When disabled, the /api/ai/chat and /api/ai/explain endpoints return HTTP 503. Note: the approval routes (GET /api/ai/approvals/[id] and POST /api/ai/approvals/[id]/decision) do not check AURIS_AI_ENABLED and will continue to respond normally even when AI is disabled.
Chat
/api/ai/chatRequires: admin:allSend a conversational message to the AI assistant. The assistant processes the full message history, optionally calls read-only tools to look up live tenant data, and returns a reply. Write actions are never executed directly — the assistant raises an approval request that must be confirmed by the caller before the action runs.
Request body
| Field | Type | Required | Description |
|---|---|---|---|
messages | Message[] | Yes | Conversation history. Min 1, max 40 items. Each content max 4 000 chars. |
mode | string | No | Assistant mode: info, plan, agent, or bypass. Default: info. |
effort | string | No | Response quality/cost level: low, medium, high, or max. Default: medium. |
reasoning | boolean | No | Override the effort-derived thinking toggle. |
context | string | No | Current page label (max 500 chars) injected into the system prompt. |
locale | string | No | BCP 47 locale tag used to localise the assistant’s reply. |
customInstructions | string | No | Additional instructions appended to the system prompt (max 4 000 chars). |
platformPolicy | object | No | Restrict the assistant’s capabilities for this call. See Platform Policy. |
approval | object | No | Approval decision for a pending action. See Approval Workflow. |
contextTags | ContextTag[] | No | Up to 20 structured context tags ({id, label, type?, value?}). |
skillId | string | No | Activate a named skill loaded by the assistant runtime. |
attachments | Attachment[] | No | Up to 10 file attachments. Each max 25 MB ({name, size, type, dataBase64}). |
questionAnswer | object | No | Reply to a clarifying question raised in the previous turn ({questionId, answer}). |
Message object
{ "role": "user", "content": "How many active sessions does [email protected] have?" }| Field | Type | Values |
|---|---|---|
role | string | user | assistant |
content | string | Message text, max 4 000 chars |
Response body
{
"reply": "Alice currently has 3 active sessions across 2 devices.",
"reasoning": "...",
"reasoningDurationMs": 1240,
"toolCalls": [
{
"id": "call_abc",
"name": "list_sessions",
"arguments": "{\"userId\":\"usr_abc123\"}",
"result": "{\"sessions\":[\"...\"]}",
"durationMs": 182
}
]
}| Field | Type | Description |
|---|---|---|
reply | string | The assistant’s text response. |
reasoning | string | Internal chain-of-thought, present only when thinking is enabled (effort high/max or reasoning: true). |
reasoningDurationMs | number | Time spent in the thinking phase (ms). |
actionRequest | object | Present when the assistant has identified a write action requiring approval. See Action Request. |
actionResult | object | Present when an approved action has been executed. Fields: actionId, actionName, status (executed|rejected|failed), details (string, optional). |
stepUp | object | Present (HTTP 428) when the action requires step-up authentication before it can proceed. See Step-Up Authentication. |
question | object | A clarifying question from the model. Fields: id, text, suggestedAnswers?, required?. Resend with questionAnswer to answer. |
toolCalls | object[] | Read-tool invocations made during this turn. arguments and result are JSON-serialised strings, not parsed objects. |
Modes
The mode field changes the assistant’s behaviour and prompt strategy.
| Mode | Behaviour |
|---|---|
info | Default. Answers questions, explains concepts, and looks up live data. Does not propose actions. |
plan | Produces a numbered execution plan with prerequisites, risk notes, and a rollback strategy. Useful before making changes. |
agent | Action-oriented. Distinguishes between read-safe operations the assistant can perform immediately and write operations that require explicit approval. |
bypass | Elevated operator mode for full-platform orchestration. Blocked unless the environment variable AURIS_AI_ALLOW_BYPASS=true is set. Returns HTTP 403 otherwise. |
Effort levels
The effort field controls the model’s token budget, temperature, and whether chain-of-thought reasoning is enabled.
| Level | Max tokens | Temperature | Reasoning effort | Thinking enabled |
|---|---|---|---|---|
low | 1 200 | 0.6 | minimal | No |
medium | 2 200 | 0.7 | minimal | No |
high | 3 200 | 0.75 | medium | Yes |
max | 4 096 | 0.8 | high | Yes |
Setting reasoning: false forces reasoning effort to minimal regardless of effort. Setting reasoning: true raises it to at least medium.
Tool loop
In agent mode the assistant has access to a set of read-only tools it can invoke autonomously (up to 5 rounds per request). Write actions are never invoked by the model directly — they are always surfaced as an approval request.
Read tools (no approval required):
search_users, get_user, get_user_roles, get_user_groups, list_audit_events, count_active_sessions, list_sessions, list_roles, get_role, list_groups, list_applications, get_application, check_url_connectivity, list_ip_rules, list_webhooks, get_security_defaults, get_tenant_info, list_security_alerts, get_mfa_stats, list_trusted_devices, list_organizations, get_organization, list_sso_connections, list_email_templates, fga_check, fga_list_tuples, list_log_streams, get_log_stream, list_pending_events
Write actions (always require approval):
user_disable, user_enable, session_revoke, session_revoke_all, user_force_password_reset, mfa_reset, ip_rule_block, ip_rule_disable, ip_rule_enable, ip_rule_delete, create_ip_allow_rule, send_password_reset_email, revoke_trusted_device, revoke_all_trusted_devices, acknowledge_security_alert, resolve_security_alert, test_log_stream, enable_log_stream, disable_log_stream, delete_log_stream, application_disable, application_enable, assign_role_to_user, remove_role_from_user, create_user, create_application, create_webhook, enable_webhook, disable_webhook, create_organization, update_organization, create_role, create_group, add_user_to_group, remove_user_from_group, delete_group, add_organization_member, remove_organization_member, update_organization_member_role, send_verification_email, update_security_defaults, update_user, update_application, regenerate_client_secret, delete_user, delete_application, delete_webhook, delete_organization, delete_role, fga_write_tuple, fga_delete_tuple
Platform policy
The optional platformPolicy field scopes down what the assistant is allowed to do for this call. This is useful when embedding the assistant in a context where only a subset of capabilities should be available.
{
"platformPolicy": {
"name": "read-only-audit",
"capabilities": ["list_audit_events", "get_user"],
"allowRead": true,
"allowWrite": false,
"permissionMode": "strict"
}
}| Field | Type | Description |
|---|---|---|
name | string | Label for the policy (used in audit logs). |
capabilities | string[] | Allowlist of tool names available to the model. Omit to allow all. |
allowRead | boolean | Whether read tools may be called. |
allowWrite | boolean | Whether write-action approval requests may be raised. |
permissionMode | string | ask — prompt user for confirmation; strict — deny any unlisted capability; none — no enforcement. |
Clarifying questions
The model may embed a fenced block in its reply to ask a clarifying question before proceeding. The server strips the block from the visible reply and surfaces it in the question field.
The model appends to its text reply:
```clarify
{"id":"q-env","text":"Which environment should I target?","suggestedAnswers":["staging","production"]}
The client should display the question to the operator and re-send the conversation with `questionAnswer: { questionId: "q-env", answer: "production" }`.
---
## Explain
<ApiEndpoint method="POST" path="/api/ai/explain" permission="admin:all">
Ask the assistant to explain an opaque object in plain language — for example a raw audit log
entry, a token exchange request, a CIBA callback, or a security event. The object is
automatically redacted of sensitive keys before being sent to the model.
</ApiEndpoint>
### Request body
```json
{
"topic": "token-exchange",
"payload": {
"grant_type": "urn:ietf:params:oauth:grant-type:token-exchange",
"subject_token_type": "urn:ietf:params:oauth:token-type:access_token",
"requested_token_type": "urn:ietf:params:oauth:token-type:refresh_token"
},
"context": { "clientId": "app_billing" },
"locale": "en"
}| Field | Type | Required | Description |
|---|---|---|---|
topic | string | No | Hint for the model: risk-assessment, ciba-request, token-exchange, device-flow, or general. Default: general. |
payload | object | Yes | The raw object to explain. |
context | object | No | Optional surrounding context (e.g., client metadata). |
locale | string | No | BCP 47 locale. |
The following keys are automatically redacted from payload before being sent to the model: token, accessToken, idToken, refreshToken, authorization, password, secret, clientSecret. Redacted values are replaced with the string [REDACTED].
Response body
{
"success": true,
"data": {
"summary": "This is a standard RFC 8693 token exchange request. The caller is trading an existing access token for a refresh token.",
"findings": [
"The subject token type is an access token, which is the standard input for this grant.",
"The requested token type is a refresh token — this extends the session lifetime."
],
"recommendations": [
"Verify that the client is authorised to perform token exchange (check the allowed_grant_types list for app_billing).",
"Confirm that the subject token has not expired before processing."
],
"confidence": "high"
}
}| Field | Type | Description |
|---|---|---|
summary | string | One-paragraph plain-language explanation. |
findings | string[] | Key observations extracted from the payload. |
recommendations | string[] | Actionable next steps. |
confidence | string | Model confidence in the explanation: low, medium, or high. |
Action Approvals
When the assistant identifies a write action in the conversation, it does not execute the action immediately. Instead it creates an AiActionApproval record and returns an actionRequest object in the chat response. The operator must explicitly approve or reject the action before it runs.
This design means the assistant can never modify tenant state without a human decision in the loop.
Approval lifecycle
pending --> approved --> executed
--> failed
--> rejected
--> expired (after 10 minutes)Approval records are retained for 30 days after expiry, then deleted.
Action risk levels
Every write action is assigned a risk level that determines whether step-up authentication is required.
| Risk | Step-up required | Example actions |
|---|---|---|
low | No | send_password_reset_email, acknowledge_security_alert, test_log_stream |
medium | No | user_enable, session_revoke, create_user, add_user_to_group |
high | No | user_disable, session_revoke_all, delete_user, delete_organization, fga_write_tuple |
critical | Yes (AAL2) | Actions classified critical require requiredAcr: aal2 step-up before execution. |
Action request object
When a write action is identified, the chat response includes an actionRequest object:
{
"actionRequest": {
"actionId": "apr_hk3mn9xq",
"actionName": "user_disable",
"summary": "Disable the account for [email protected] (usr_abc123).",
"risk": "high",
"requiresApproval": true,
"requiresStepUp": false,
"requiredAcr": null,
"expiresAt": "2025-02-18T10:10:00Z",
"approvalToken": "eyJ..."
}
}To approve the action inline, resend the conversation with the approval field:
{
"messages": [...],
"approval": {
"actionId": "apr_hk3mn9xq",
"approvalToken": "eyJ...",
"approved": true,
"decisionReason": "Confirmed with the account owner."
}
}approvalToken is optional in the inline approval object.
Alternatively, use the dedicated approval endpoints below.
Get Approval
/api/ai/approvals/[id]Requires: admin:allRetrieve the current state of a pending or completed approval record.
Success response
{
"actionId": "apr_hk3mn9xq",
"actionName": "user_disable",
"summary": "Disable the account for [email protected].",
"risk": "high",
"status": "pending",
"requiresStepUp": false,
"requiredAcr": null,
"expiresAt": "2025-02-18T10:10:00Z",
"createdAt": "2025-02-18T10:00:00Z",
"approvedAt": null,
"rejectedAt": null,
"executedAt": null,
"decisionReason": null,
"actionResultStatus": null,
"actionResultDetails": null
}The response is the ActionApprovalView object returned directly — there is no ok/data envelope.
Error response (404)
{ "error": "Approval request not found" }Error codes
| HTTP | Description |
|---|---|
| 404 | No approval record with this ID. |
Submit Decision
/api/ai/approvals/[id]/decisionRequires: admin:allApprove or reject a pending action. If the action requires step-up authentication, the server returns HTTP 428 with a challenge. Complete the challenge and re-submit with the step-up response to proceed.
Request body
{
"approved": true,
"decisionReason": "Verified with the account owner before proceeding.",
"stepUp": {
"challengeId": "chal_xyz789",
"method": "totp",
"response": "482910"
}
}| Field | Type | Required | Description |
|---|---|---|---|
approved | boolean | Yes | true to approve, false to reject. |
decisionReason | string | No | Freeform reason logged to the audit trail (max 2 000 chars). |
stepUp | object | Conditional | Step-up response. Required if the server previously returned HTTP 428 for this approval. |
Success response
{
"status": "executed",
"approval": { "actionId": "apr_hk3mn9xq", "actionName": "user_disable", "...": "ActionApprovalView fields" },
"actionResult": {
"actionId": "apr_hk3mn9xq",
"actionName": "user_disable",
"status": "executed",
"details": "User usr_abc123 disabled successfully."
}
}The response is the ApprovalDecisionResult object returned directly — there is no ok/data envelope. actionResult.details is a plain string (or absent), not a structured object.
The status field in the response reflects the outcome:
| Status | Meaning |
|---|---|
approved | Decision recorded; action is queued for execution (async). |
rejected | Decision recorded; action will not run. |
executed | Action ran successfully in this call. |
failed | Action was approved but execution failed. Check actionResult.details. |
step_up_required | HTTP 428. Step-up challenge issued. Re-submit with stepUp fields. |
Error codes
| HTTP | Description | Response body |
|---|---|---|
| 404 | Approval record does not exist. | { "error": "..." } |
| 410 | The 10-minute TTL has elapsed. The action cannot be executed. | { "error": "..." } |
| 403 | Step-up challenge response was invalid or expired. | { "error": "<message>" } |
Step-Up Authentication
Critical-risk actions require the operator to complete an MFA challenge before the action executes. This is separate from the session-level MFA and provides a fresh proof of presence.
Flow
- The operator calls
POST /api/ai/approvals/[id]/decisionwithapproved: truebut withoutstepUpfields. - If the action requires step-up (
requiresStepUp: true), the server returns HTTP 428:
{
"status": "step_up_required",
"approval": { "actionId": "apr_hk3mn9xq", "...": "ActionApprovalView fields" },
"stepUp": {
"challengeId": "chal_xyz789",
"availableMethods": ["totp", "webauthn"],
"expiresIn": 300,
"requiredAcr": "aal2"
}
}There is no ok/error envelope — the body is the ApprovalDecisionResult object with status: "step_up_required" directly.
- The client presents the MFA challenge to the operator.
- The operator completes the challenge. The client re-submits the decision with the
stepUpobject:
{
"approved": true,
"stepUp": {
"challengeId": "chal_xyz789",
"method": "totp",
"response": "482910"
}
}- The server validates the response and, if successful, executes the action.
ACR level hierarchy
| Level | Numeric value | Description |
|---|---|---|
aal1 | 1 | Password only |
aal2 | 2 | Password + second factor (TOTP, WebAuthn, SMS) |
aal3 | 3 | Password + hardware-bound second factor |
The operator’s session ACR must meet or exceed requiredAcr after step-up.
Audit Events
Every state transition in the approval lifecycle emits an audit event visible in the Audit Logs API.
| Event | Trigger |
|---|---|
AI_ACTION_APPROVAL_REQUESTED | An approval record was created. |
AI_ACTION_STEP_UP_REQUIRED | HTTP 428 was returned; challenge issued. |
AI_ACTION_STEP_UP_VERIFIED | Step-up challenge passed. |
AI_ACTION_APPROVED | Operator approved the action. |
AI_ACTION_REJECTED | Operator rejected the action. |
AI_ACTION_EXECUTED | Action ran successfully. |
AI_ACTION_FAILED | Action was approved but execution failed. |
AI_ACTION_APPROVAL_EXPIRED | Approval record reached its 10-minute TTL without a decision. |
Permissions
All endpoints in this group require the admin:all permission. There is no read-only scoping — the assistant’s ability to query live tenant data and raise approval requests for write actions is considered an administrative capability.
The bypass mode (mode: "bypass") requires the AURIS_AI_ALLOW_BYPASS=true environment variable in addition to the admin:all permission. Requests with mode: "bypass" return HTTP 403 unless this flag is explicitly set.
Environment Variables
| Variable | Default | Description |
|---|---|---|
AURIS_AI_ENABLED | true | Set to false to disable AI on /api/ai/chat and /api/ai/explain (returns 503). The approvals routes are unaffected. |
AURIS_AI_API_KEY | — | Provider API key. If absent, /api/ai/chat and /api/ai/explain return HTTP 503 with code AI_DISABLED — same effect as AURIS_AI_ENABLED=false. Also consumed by the LiteLLM gateway. |
AURIS_AI_MODEL | MiniMax-M2.7 | Model identifier forwarded to the LiteLLM gateway. |
AURIS_AI_API_URL | https://api.minimax.io/v1/text/chatcompletion_v2 | Provider base URL. Set to http://litellm:4000/v1/chat/completions to route through the internal LiteLLM proxy. |
AURIS_AI_TIMEOUT_MS | 30000 | Per-request timeout in milliseconds. Exceeded requests return HTTP 504. |
AURIS_AI_SYSTEM_INSTRUCTIONS | — | Additional instructions appended to the system prompt of every chat request. |
AURIS_AI_ALLOW_BYPASS | false | Must be true to permit mode: "bypass" requests. |
Related
- Audit Logs API — Query
AI_ACTION_*events - Sessions API —
count_active_sessionsandsession_revoketools used by the assistant - Security API —
list_security_alertsand IP rule tools used by the assistant - Fine-Grained Authorization API —
fga_checkandfga_write_tupletools used by the assistant