How It Works
Point
Set the OpenAI client’s
base_url to your Praxis Chat Completions endpoint.Authenticate
Pass your Praxis JWT via the
x-access-token header to identify the user session.Chat
Send OpenAI-format messages. Pria routes them to your Digital Twin — the one your token belongs to, or the one you name in the
x-praxis-institution-public-id header (institution is the API’s word for a Digital Twin).Prerequisites
Before you begin, make sure you have:- A Praxis AI account with access to at least one Digital Twin
- The Digital Twin’s Public ID (a UUID like
e455529a-4f51-479e-94fc-bbebb41d19a1) — found in the Digital Twin’s administration panel - A valid Praxis JWT token (
x-access-token) — obtained when a user authenticates with Praxis (see Authentication) - Chat Completions enabled on the Digital Twin — this integration is off by default; an administrator must turn it on for the Digital Twin
The Chat Completions endpoint is part of the Pria platform itself — the base URL is your Pria server’s
/api/ai path (e.g. https://pria.praxislxp.com/api/ai). It is disabled per Digital Twin by default; if requests return 403 chat_completion_disabled, ask the Digital Twin’s administrator to enable the Chat Completions endpoint in its configuration.Quick Start
1
Install the OpenAI SDK
2
Configure the client
Point the SDK to your Praxis endpoint and pass your authentication token.
The shortest version: put the JWT in
api_key and drop the header. The SDK sends api_key as Authorization: Bearer …, which Praxis accepts, so OpenAI(api_key=your_jwt, base_url=…) is enough on its own. The x-access-token header above is shown because it’s the form every other Pria example uses; either works, and you never need both.What you cannot do is pass a personal pria_… API key here — exchange it for a JWT first (see Authentication below).3
Send a message (streaming)
Pass your Digital Twin Public ID as the
model parameter and set stream=True. The model value does not choose anything — Pria works out which Twin to use from your token (or from the x-praxis-institution-public-id header); the value is only echoed back in the response.The endpoint always streams — responses are delivered as OpenAI-format SSE chunks regardless of the
stream flag, so use your SDK’s streaming mode.The
model value is informational (it is echoed back in the SSE chunks): the request is served by the Digital Twin resolved from your token — the user’s current Digital Twin by default. If your user belongs to several Digital Twins, target one explicitly with the x-praxis-institution-public-id header (see Context Headers).Authentication
Every request carries a Praxis JWT — the token that identifies the user and authorizes access to their Digital Twins. You can present it in either of two ways, whichever your client makes easier:api_key, which is why passing the JWT as the api_key works with no extra configuration.
Getting a JWT from a personal API key (server-to-server)
For scripts and server-to-server integrations, exchange a personal API key (prefixedpria_) for a JWT, then pass that JWT in the x-access-token header. The raw pria_… key is not accepted directly — exchange it first:
Context Headers
Optional headers let you pass conversation metadata to the Digital Twin. These enrich the interaction context without affecting authentication.Context headers are useful when your application manages multiple conversations or needs to target a specific assistant within a Digital Twin.
Message Roles
The API accepts standard OpenAI message roles with the following behavior:Send the turns in order, alternating. Pria tidies up what you send — consecutive messages from the same speaker are joined, and the history is trimmed so it reads as a proper back-and-forth ending on the user — because some providers reject anything else. You’ll get a sensible answer either way, but a clean alternating array is what you’ll see reflected back.
You can send only the current user message (Pria tracks the conversation via
x-praxis-conversation-id), or pass your own running message array — prior user/assistant turns you include are replayed as history for the active turn.Response Format
The endpoint always streams. Responses arrive as standard OpenAI SSE chunks:choices[0].delta.content.
Supported Parameters
The Digital Twin’s administrator controls this endpoint, under Chat Completion in the twin’s settings. Turning on Enable Chat Completion endpoint reveals three overrides that apply to inbound requests only, leaving the twin’s in-app behaviour untouched:
- Chat Completion Model — a different model for API traffic
- Chat Completion Max Output Tokens — a ceiling on the length of each answer
- Chat Completion Reasoning Effort — leave on Inherit, or pick a level; None is the usual choice for voice agents, where thinking time is silence
Error Handling
Two shapes, depending on how far the request got. Errors raised by the endpoint itself use the OpenAI envelope your SDK already understands:Common Errors
A missing token is 403, not 401. It’s an easy one to trip over when your retry logic keys on 401 — a client that only refreshes on 401 will loop on 403 forever without ever sending a token.
Multi-Provider Routing
The Chat Completions API is a front door to the Pria platform, not a thin proxy to a single model provider. Behind the URL, Pria selects the underlying provider and model based on the Digital Twin’s configuration. As of today, Pria can route to:
The
model parameter you pass in the request is not a provider model ID — it does not select the model at all. The effective model follows the Digital Twin’s configuration cascade (assistant override → the twin’s Chat Completions model override → the twin’s conversation model). If the Twin’s admin changes the underlying model from Claude to GPT‑5, your code does not change. When the twin has a Fallback conversation model and its model fails before it has written anything, the fallback answers instead of an error.
Per‑Provider Behavioural Differences
Because Pria forwards to many providers, some advanced behaviours are provider‑dependent and respect the Twin’s configuration rather than the request payload:- Reasoning effort — honored by reasoning‑capable models (OpenAI GPT‑5 family and o‑series, xAI reasoning variants); models that always reason ignore the setting. Configured per Twin, with a Chat‑Completions‑specific override available.
- Thinking tokens — Anthropic Claude and Google Gemini thinking models support extended thinking budgets, configured per Twin.
- Image generation — supported through OpenAI (GPT Image 2.5), Google (Gemini image models), xAI (Grok Imagine), and Stability AI (Stable Image).
- Prompt caching — applied automatically where the underlying provider supports it; cached reads are billed at a discount.
- Tool calls — the Digital Twin’s server‑side tools (RAG, web search, charts, connectors, MCP) run automatically. Client‑supplied
tools/tool_choiceare not forwarded (see Supported Parameters).
Provider Authentication Errors
When the Twin is configured to use a provider that requires its own credentials (BYOT — Bring Your Own Tokens), errors from the underlying provider are surfaced back to you as standard OpenAI‑style errors:
See BYOT (Bring Your Own Tokens) for how Twin admins configure provider keys.
Cost & Credits
Chat completions consume Pria credits, billed by token usage and the underlying provider’s price tier. Usage is metered server-side per turn — the SSE stream does not include an OpenAI-styleusage block, so review consumption in Pria’s credit history rather than in the response payload.
- Cached prompt tokens (when the provider supports caching) are billed at a discounted rate.
- Streaming is the only mode, and it is billed like any conversation turn.
- Embedded RAG retrieval runs as part of the Digital Twin’s response and is included in the credit cost — you do not pay separately for vector search.
Chat Completions vs. the Pria Runtime API
Pria exposes two complementary APIs. They look similar but behave very differently — choose based on whether you want stateless OpenAI‑style requests or full Pria session semantics.
Rule of thumb: if your code already uses the OpenAI SDK and you want to point it at Pria with minimal changes, use the Chat Completions API. If you’re building a new client from scratch and want access to every Pria capability (per‑message tool events, citations, KAG augmentations, structured memory updates), use the Runtime API.
See the API Reference for both.
Related
- API Reference — Full REST API documentation with streaming details
- AI Models — Provider catalog, reasoning effort, thinking, and image generation behaviour per provider
- API Keys — Issue and rotate the API keys used to obtain Praxis JWTs
- BYOT (Bring Your Own Tokens) — How Twin admins configure per‑provider credentials
- Plans & Credits — How token usage maps to credits
- MCP Server — Connect Pria to custom LLM workflows
- Web SDK — Embed the full Digital Twin UI in your web app
- JavaScript SDK — Programmatic control of the Pria interface