Skip to content

Plain chat completions

Call the gateway's chat completions directly, without an agent, with an API token.

An agent is usually what you want: the server owns its instructions, model and tool contract, and the SDK runs its tools. Without an agent, the client calls the plain /v1/chat/completions route. You pick the model, write the messages, and run any tool calls yourself. The server's policies still apply, and a refusal arrives as part of the answer. See Chat completions for the API itself.

The client without an agent

Create a FadenClient with an API token and a model:

import { FadenClient, StaticTokenProvider } from "@fadenstack/client";

const client = new FadenClient({
  serverUrl: "https://ai.example.internal",
  credentials: new StaticTokenProvider(process.env.FADEN_API_TOKEN ?? ""),
  model: "my-model", // a model the server offers
});

StaticTokenProvider sends the same token with every request. It throws a TypeError for an empty token.

Option What it is
serverUrl The Fadenstack server. Required.
credentials Where access tokens come from: a StaticTokenProvider with an API token, or an agent sign-in.
model The model for the plain routes. An agent's model is the agent's, whatever this says.
agent The agent's slug or id. With it, chat goes to the agent's route instead; FadenClient.forAgent sets it.
toolMode The tool mode sent with each request, unless a request says otherwise. Default ask.
gatewayEvents Whether to ask the gateway for its events: tool rounds, approvals, notices. Default on.
apiUrl The server's backend for sign-in, when it is not at serverUrl.
fetch The fetch to use. The global one by default.
timeoutMs How long one request may take, in milliseconds. The default is five minutes. A streamed answer that has started is not cut off by it.
allowInsecureHttp Allows an http:// address. Only for a development server on your own machine.

client.getFeatures() returns the server's edition and features (GET /v1/features, on a server from the next release, after 0.4.1).

A whole answer

client.chat.complete(messages, options?) returns the whole answer at once:

const response = await client.chat.complete([
  { role: "system", content: "Answer in one sentence." },
  { role: "user", content: "What does a heat exchanger do?" },
]);

if (response.refused) console.log(`Refused: ${response.refused.message}`);
else console.log(response.text);

The ChatResponse:

Field What it is
text The answer.
reasoning The model's reasoning, for models that report it. Empty otherwise.
toolCalls The tool calls the model made.
finishReason stop, length, tool_calls, content_filter, or another value the model gives.
usage promptTokens, completionTokens, totalTokens, and cachedPromptTokens when the server reports it.
events The gateway's events for the request.
refused Set when the gateway refused the prompt or stopped the answer: a notice with the outcome and the message to show.
id, model The response's id and the model that answered.

A refusal is an answer, not an error: refused is set and finishReason is content_filter. Only real failures throw. See Outcomes and errors.

A streamed answer

client.chat.stream(messages, options?) yields the answer as it streams:

for await (const update of client.chat.stream(messages, { temperature: 0.2 })) {
  switch (update.type) {
    case "text": process.stdout.write(update.text); break;
    case "refused": console.log(`\nRefused: ${update.notice.message}`); break;
    case "finish": console.log(`\n(${update.reason})`); break;
  }
}
type Fields What it is
text text A piece of the answer.
reasoning text A piece of the model's reasoning.
tool_calls calls The tool calls, whole, once the model has written them.
usage usage The tokens used.
event event One of the gateway's events.
refused notice The gateway refused the prompt or stopped the answer.
finish reason, responseId, model The answer ended.

Request options

Both methods take ChatOptions:

Option What it is
model The model for this request, instead of the client's.
temperature, topP, maxTokens, seed, stop Sent as the request fields temperature, top_p, max_tokens, seed and stop.
reasoningEffort How much a reasoning model thinks first: none, minimal, low, medium, high or xhigh. none answers at once.
responseFormat { type: "json_object" }, or { type: "json_schema", name?, schema } for an answer in a given shape.
tools Tools the model may call: { name, description?, parameters? }, with parameters as JSON Schema.
toolChoice auto, none, required, or { name } for one tool.
toolMode The tool mode for this request.
conversationId Groups usage per conversation, without the server keeping the conversation.
turnId An id for the turn.
sessionId A server session, for the plain routes.
skills Skills picked by slug for this message.
extra More request fields, sent as they are.
signal Stops the request.

Tool calls

The client gives you the model's tool calls whole: each ToolCall has an id, a name, the arguments as the model wrote them (argumentsJson), and the parsed arguments, or argumentsError when they are not a JSON object. Running the tools and sending their results back is up to you: add the model's message with its tool_calls, a tool message per result, and call the model again.

An agent does all of this for you, with tool classes, tool modes and approvals. See Build an agent.

Other routes

client.transport makes authenticated requests to other routes, with the same tokens and typed errors:

const features = await client.transport.get<{ edition: string; features: string[] }>(
  client.transport.gateway("v1/features"),
);

gateway(path) builds an address under the server, api(path) one under the backend. get, post, patch, delete and json send JSON and throw a FadenApiError for an error status.