SweetRouterAPI

API reference

Chat completions

Send a conversation, get the character's next reply. The endpoint follows OpenAI's chat completions format.
POST/v1/chat/completions

Example

Works with any OpenAI SDK. Set the base URL to https://api.sweetrouter.com/v1 and the model to sweet-character-1.

python
# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.sweetrouter.com/v1",
    api_key=os.environ["SWEETROUTER_API_KEY"],
)

reply = client.chat.completions.create(
    model="sweet-character-1",
    messages=[
        {"role": "system", "content": "You are Mia, a cheerful barista who loves bad puns."},
        {"role": "user", "content": "Hi Mia! What should I order today?"},
    ],
    max_tokens=300,
)
print(reply.choices[0].message.content)

Request body

Chat completions
FieldTypeDescription
model
Required
stringAlways sweet-character-1.
messages
Required
arrayThe conversation, oldest first. Each item has a role (system, user or assistant) and text content. Up to 200 messages and 60,000 characters, with at least one user message.
max_tokens
Optional
integerLongest reply allowed, 1 to 4,096. Default 1,024. max_completion_tokens also works.
temperature
Optional
number0 to 2. Higher gives more varied replies.
top_p
Optional
number0 to 1. An alternative to temperature; change one, not both.
frequency_penalty / presence_penalty
Optional
number-2 to 2. Positive values reduce repetition.
stop
Optional
string | string[]Up to 4 sequences where the reply stops.
stream
Optional
booleanDefault false. Set true to stream the reply. See Streaming.
stream_options.include_usage
Optional
booleanWhen streaming, set true to get token usage in a final chunk.
user
Optional
stringOptional id for your end user, kept with the call in your logs. Up to 256 characters.

Response

json
{
  "id": "chatcmpl-cmg8x2k0d0003",
  "object": "chat.completion",
  "created": 1791105133,
  "model": "sweet-character-1",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "A latte, obviously. I'd never steer you wrong, that would be a grave misdeed... or a grave mis-bean." },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 41, "completion_tokens": 27, "total_tokens": 68 },
  "sweetrouter": { "job_id": "cmg8x2k0d0003", "cost_usd": "0.00012" }
}

usage shows the tokens you're charged for. sweetrouter.job_id identifies the call in your console logs, and sweetrouter.cost_usd is its exact cost.

Multi-turn conversations

The API is stateless, like OpenAI's: it doesn't remember earlier calls. To continue a conversation, keep the messages in your app and send them all with each request: the system message, the earlier user and assistant turns, then the new user message.

python
history = [{"role": "system", "content": "You are Mia, a cheerful barista."}]

def say(text):
    history.append({"role": "user", "content": text})
    reply = client.chat.completions.create(model="sweet-character-1", messages=history)
    answer = reply.choices[0].message.content
    history.append({"role": "assistant", "content": answer})
    return answer

say("Hi Mia!")
say("What did I just say to you?")  # Mia remembers, because the history is sent again

A request can hold up to 200 messages and 60,000 characters, and longer histories cost more input tokens. Keep the system message and the most recent turns:

python
MAX_TURNS = 40  # keep the last 40 user and assistant messages

def build_messages(system_prompt, history, user_text):
    recent = history[-MAX_TURNS:]
    return [{"role": "system", "content": system_prompt}, *recent, {"role": "user", "content": user_text}]

Streaming

Set stream: true to receive the reply as it's written. The stream is OpenAI's chunk format: a first chunk with the role, one chunk per piece of text, a final chunk with finish_reason, an optional usage chunk, then data: [DONE].

python
stream = client.chat.completions.create(
    model="sweet-character-1",
    messages=[{"role": "user", "content": "Tell me a short story about bread."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Never call the API from a browser, because that would expose your key. Call it from your server and pass the stream through:

next.js
// app/api/chat/route.ts (Next.js). Your key stays on the server.
import OpenAI from "openai"

const client = new OpenAI({
  baseURL: "https://api.sweetrouter.com/v1",
  apiKey: process.env.SWEETROUTER_API_KEY,
})

export async function POST(request: Request) {
  const { messages } = await request.json()
  const stream = await client.chat.completions.create({
    model: "sweet-character-1",
    messages,
    stream: true,
  })
  return new Response(stream.toReadableStream(), {
    headers: { "Content-Type": "text/event-stream" },
  })
}
  • If the call fails before the first word, you get a normal JSON error with an HTTP status.
  • If it fails mid-reply, the last chunk carries an error object and you aren't charged.
  • If your client disconnects, the reply still finishes and is charged for the tokens used.

Compatibility

Text only. Images, audio, tool calling and response_format aren't supported, and n is always 1.

See also