Chat completions¶
POST /v1/chat/completions works like OpenAI's, with a few additions for skills, tools and the organisation's rules.
Request¶
model and messages are required. model is a model name from /v1/models, chosen by your administrator.
| Fields | Treatment |
|---|---|
messages, stream, temperature, top_p, max_tokens, stop, presence_penalty, frequency_penalty, logit_bias, seed, n, user, tools, tool_choice, response_format, reasoning_effort |
Passed to the model, as in OpenAI's API |
max_completion_tokens, parallel_tool_calls, logprobs, stream_options |
Ignored. Streamed answers always end with a usage chunk. |
skills |
Fadenstack: skills to use for this request. See below. |
session_id |
Fadenstack: keep the conversation on the server, in that chat session. |
Requests are limited to 2 MiB.
curl -s https://ai.example.internal/v1/chat/completions \
-H "Authorization: Bearer $FADEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "team-assistant",
"skills": ["incident-report"],
"messages": [{"role": "user", "content": "Write up last night'\''s outage of line 3."}]
}'
Skills¶
A skill is a set of instructions your organisation shares, for example a report format. Name skills in a request in either way:
- the
skillsfield, with their names:"skills": ["incident-report"] - in the message, with a dollar sign:
Use $incident-report for last night's outage.
You can use the skills you may see in the console. See Skills.
Tools¶
Your tools work as in OpenAI's API: define them in tools, and when the model calls one you get a tool_calls message back, run it, and send the result.
Tools on the server. When the model can call tools, the server may also offer it tools of its own, such as searching your organisation's documents or finding an MCP tool that fits. The server runs those itself, between the model's steps; you get the final answer. Set "tool_choice": "none" to switch this off for a request.
Tool mode¶
Send the session's tool mode in a header, and the server applies it to its own tools and to yours:
off, read_only, ask (the default) or auto. In off and read_only, tools without a class are hidden. See Tools and approvals.
Server version
Tool modes, rules and the extra events need a Fadenstack server from the next release, after 0.4.1.
When a rule refuses¶
The organisation's rules can refuse a request or an answer:
- A refused request gets status 422 with code
policy_blocked. Nothing was sent to a model. Thefaden.governanceobject says which stage refused it and carries a message for the user. - A refused answer ends with
finish_reason: "content_filter". The user's message was processed, but the answer is withheld. - A refused tool call is reported to the model as refused, and the answer goes on.
Show the message to the user as it is. The SDKs turn these into outcomes.
Extra events¶
With X-Faden-Events: ag-ui, a streamed answer also carries events about server tools, sources, personal data and refusals. See Events.