Developer Guide
Getting started with Ziwol API
Integrate your application with Sava AI through Ziwol API, or install Sava Omni Agent to run development tasks directly from your terminal.
Available Models
Choose a model based on output requirements, context size, account role, and whether agent capabilities are required.
| Model ID | Name | Max Output | Context | Minimum Role |
|---|---|---|---|---|
sava-flash |
Sava Flash | 5,000 tokens | 32K | Free |
sava-2.0 |
Sava 2.0 | 8,000 tokens | 90K | Free |
sava-omni |
Sava Omni | 8,000 tokens | 200K | Spectrum |
sava-omni-agent |
Sava Omni Agent | 64,000 tokens | 200K | Spectrum |
Sava Omni Agent Installation
Sava Omni Agent executes development tasks directly from your terminal. Installation requires a supported account role and one shell command.
Run the installer in your terminal
curl -fsSL https://api.ziwol.com/install | bash
Start Sava Omni Agent
sava
API Integration
Send requests to https://api.ziwol.com/v1 and provide your API key in the x-api-key request header.
Python
import requests
API_KEY = "sk-zw-your-api-key"
BASE_URL = "https://api.ziwol.com/v1"
payload = {
"model": "sava-flash",
"messages": [
{
"role": "user",
"content": "Hello!",
}
],
}
response = requests.post(
f"{BASE_URL}/messages",
headers={
"Content-Type": "application/json",
"x-api-key": API_KEY,
},
json=payload,
timeout=60,
)
data = response.json()
if response.ok and "content" in data:
print(data["content"][0]["text"])
else:
print(f"Error: {response.status_code}", data)
Node.js
const API_KEY = "sk-zw-your-api-key";
const BASE_URL = "https://api.ziwol.com/v1";
const response = await fetch(`${BASE_URL}/messages`, {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-api-key": API_KEY,
},
body: JSON.stringify({
model: "sava-flash",
messages: [
{
role: "user",
content: "Hello!",
},
],
}),
});
const data = await response.json();
if (!response.ok) {
throw new Error(`Request failed: ${response.status}`);
}
console.log(data.content[0].text);
cURL
curl https://api.ziwol.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: sk-zw-your-api-key" \
-d '{
"model": "sava-flash",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'
Streaming with Server-Sent Events
Set "stream": true to receive response events as they are generated.
curl https://api.ziwol.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: sk-zw-your-api-key" \
-d '{
"model": "sava-2.0",
"stream": true,
"messages": [
{
"role": "user",
"content": "Tell me a story"
}
]
}'
Billing Modes
Each API key can use one of three billing modes. Select the mode that best matches your budget and availability requirements.
Hybrid
Uses daily subscription quota first. When plan quota is exhausted, requests automatically continue using your PAYG balance.
Plan Only
Uses daily plan quota only. Requests stop when the available plan limit is reached, protecting your PAYG balance.
PAYG
Charges usage directly to your dollar balance. Daily plan token limits do not apply while sufficient balance remains.
Rate Limits and Usage Protection
Ziwol API enforces account, model, concurrency, and global usage limits to maintain predictable availability and fair usage.
Free User Quotas
Free users receive a one-time 25,000-token bonus and a daily quota of 10,000 tokens. Maximum output is limited to 5,000 tokens per request.
Global Concurrency Limit
A maximum of 2500 free users can make requests concurrently. Requests above that threshold can receive a 503 Overloaded response.
Global Token Quota
System can configure a global daily token quota. Free requests return 503 Overloaded after this quota is exhausted and resume after reset.
Automatic Maximum Tokens
If max_tokens is omitted, the service automatically applies the maximum output limit allowed by your current role and plan.
Plan Limit Enforcement
Requests with a max_tokens value above your plan allowance are rejected with 403 Forbidden. Lower the value or upgrade the plan.
Input Context Limit
Input larger than the selected model context window is rejected with 400 Bad Request. Select a larger model or reduce the submitted context.