Ziwol API
Documentation

Developer Guide

Getting started with Ziwol API

Integrate your application with Sava AI through Ziwol API, or install Sava Omni Agent to run development tasks directly from your terminal.

Available Models

Choose a model based on output requirements, context size, account role, and whether agent capabilities are required.

Model ID Name Max Output Context Minimum Role
sava-flash Sava Flash 5,000 tokens 32K Free
sava-2.0 Sava 2.0 8,000 tokens 90K Free
sava-omni Sava Omni 8,000 tokens 200K Spectrum
sava-omni-agent Sava Omni Agent 64,000 tokens 200K Spectrum

Sava Omni Agent Installation

Sava Omni Agent executes development tasks directly from your terminal. Installation requires a supported account role and one shell command.

1

Run the installer in your terminal

Shell
curl -fsSL https://api.ziwol.com/install | bash
2

Start Sava Omni Agent

Shell
sava
Restricted access. Sava Omni Agent is available to Spectrum, Business.

API Integration

Send requests to https://api.ziwol.com/v1 and provide your API key in the x-api-key request header.

Keep API keys on your server. Do not expose them in client-side JavaScript, mobile application bundles, public repositories, or browser storage.

Python

Python
import requests

API_KEY = "sk-zw-your-api-key"
BASE_URL = "https://api.ziwol.com/v1"

payload = {
    "model": "sava-flash",
    "messages": [
        {
            "role": "user",
            "content": "Hello!",
        }
    ],
}

response = requests.post(
    f"{BASE_URL}/messages",
    headers={
        "Content-Type": "application/json",
        "x-api-key": API_KEY,
    },
    json=payload,
    timeout=60,
)

data = response.json()

if response.ok and "content" in data:
    print(data["content"][0]["text"])
else:
    print(f"Error: {response.status_code}", data)

Node.js

JavaScript
const API_KEY = "sk-zw-your-api-key";
const BASE_URL = "https://api.ziwol.com/v1";

const response = await fetch(`${BASE_URL}/messages`, {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "x-api-key": API_KEY,
  },
  body: JSON.stringify({
    model: "sava-flash",
    messages: [
      {
        role: "user",
        content: "Hello!",
      },
    ],
  }),
});

const data = await response.json();

if (!response.ok) {
  throw new Error(`Request failed: ${response.status}`);
}

console.log(data.content[0].text);

cURL

Shell
curl https://api.ziwol.com/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: sk-zw-your-api-key" \
  -d '{
    "model": "sava-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }'

Streaming with Server-Sent Events

Set "stream": true to receive response events as they are generated.

Shell
curl https://api.ziwol.com/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: sk-zw-your-api-key" \
  -d '{
    "model": "sava-2.0",
    "stream": true,
    "messages": [
      {
        "role": "user",
        "content": "Tell me a story"
      }
    ]
  }'

Billing Modes

Each API key can use one of three billing modes. Select the mode that best matches your budget and availability requirements.

Hybrid

Uses daily subscription quota first. When plan quota is exhausted, requests automatically continue using your PAYG balance.

Plan Only

Uses daily plan quota only. Requests stop when the available plan limit is reached, protecting your PAYG balance.

PAYG

Charges usage directly to your dollar balance. Daily plan token limits do not apply while sufficient balance remains.

Per-key tracking. Request count, input and output tokens, costs, billing mode, and last-used time are available on the Dashboard.

Rate Limits and Usage Protection

Ziwol API enforces account, model, concurrency, and global usage limits to maintain predictable availability and fair usage.

Free User Quotas

Free users receive a one-time 25,000-token bonus and a daily quota of 10,000 tokens. Maximum output is limited to 5,000 tokens per request.

Global Concurrency Limit

A maximum of 2500 free users can make requests concurrently. Requests above that threshold can receive a 503 Overloaded response.

Global Token Quota

System can configure a global daily token quota. Free requests return 503 Overloaded after this quota is exhausted and resume after reset.

Automatic Maximum Tokens

If max_tokens is omitted, the service automatically applies the maximum output limit allowed by your current role and plan.

Plan Limit Enforcement

Requests with a max_tokens value above your plan allowance are rejected with 403 Forbidden. Lower the value or upgrade the plan.

Input Context Limit

Input larger than the selected model context window is rejected with 400 Bad Request. Select a larger model or reduce the submitted context.