Skip to content

Chat completions

POST /v1/chat/completions works like OpenAI's, with a few additions for skills, tools and the organisation's rules.

Request

model and messages are required. model is a model name from /v1/models, chosen by your administrator.

Fields Treatment
messages, stream, temperature, top_p, max_tokens, stop, presence_penalty, frequency_penalty, logit_bias, seed, n, user, tools, tool_choice, response_format, reasoning_effort Passed to the model, as in OpenAI's API
max_completion_tokens, parallel_tool_calls, logprobs, stream_options Ignored. Streamed answers always end with a usage chunk.
skills Fadenstack: skills to use for this request. See below.
session_id Fadenstack: keep the conversation on the server, in that chat session.

Requests are limited to 2 MiB.

curl -s https://ai.example.internal/v1/chat/completions \
  -H "Authorization: Bearer $FADEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "team-assistant",
    "skills": ["incident-report"],
    "messages": [{"role": "user", "content": "Write up last night'\''s outage of line 3."}]
  }'

Skills

A skill is a set of instructions your organisation shares, for example a report format. Name skills in a request in either way:

  • the skills field, with their names: "skills": ["incident-report"]
  • in the message, with a dollar sign: Use $incident-report for last night's outage.

You can use the skills you may see in the console. See Skills.

Tools

Your tools work as in OpenAI's API: define them in tools, and when the model calls one you get a tool_calls message back, run it, and send the result.

Tools on the server. When the model can call tools, the server may also offer it tools of its own, such as searching your organisation's documents or finding an MCP tool that fits. The server runs those itself, between the model's steps; you get the final answer. Set "tool_choice": "none" to switch this off for a request.

Tool mode

Send the session's tool mode in a header, and the server applies it to its own tools and to yours:

X-Faden-Tool-Mode: ask

off, read_only, ask (the default) or auto. In off and read_only, tools without a class are hidden. See Tools and approvals.

Server version

Tool modes, rules and the extra events need a Fadenstack server from the next release, after 0.4.1.

When a rule refuses

The organisation's rules can refuse a request or an answer:

  • A refused request gets status 422 with code policy_blocked. Nothing was sent to a model. The faden.governance object says which stage refused it and carries a message for the user.
  • A refused answer ends with finish_reason: "content_filter". The user's message was processed, but the answer is withheld.
  • A refused tool call is reported to the model as refused, and the answer goes on.

Show the message to the user as it is. The SDKs turn these into outcomes.

Extra events

With X-Faden-Events: ag-ui, a streamed answer also carries events about server tools, sources, personal data and refusals. See Events.