DEVELOPER REFERENCE
Dedicated LLMs API
OpenAI-compatible text chat with automatic failover.
Quickstart
- Base URL
- API key
abhiaikey
POST /v1/chat/completions accepts text system, user and assistant messages. Options: stream, max_tokens (1-8192), temperature (0-2), top_p (0-1), and stop. Only n=1 is supported. Tools, images and structured output are not supported.
Models
These are custom assistant aliases, not verified releases of the similarly named vendor models. The response model field stays on your selected alias while inference may use a different configured provider.
| API model ID | Assistant name |
|---|---|
claud-opus-5 | claud opus 5 |
gpt-6.1-sol | gpt 6.1 sol |
glm-5.2 | glm 5.2 |
GET /v1/models returns exactly these three aliases. All /v1 requests require Authorization: Bearer abhiaikey or X-API-Key: abhiaikey.
Streaming
Set "stream": true to receive real server-sent events ending in data: [DONE]. Provider failover is allowed before the first output token. After output begins, interruption is reported as an SSE error instead of mixing answers from different models.
data: {"object":"chat.completion.chunk","model":"gpt-6.1-sol","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: [DONE]Errors & limits
| Status | Meaning |
|---|---|
| 400 / 413 / 415 | Invalid, oversized or unsupported request |
| 401 | Invalid Dedicated LLMs key |
| 404 | Unknown model alias or endpoint |
| 429 | Temporary rate limit |
| 503 | Service temporarily unavailable |
Requests scale across the service network, but processing capacity and upstream quotas still apply. Free inference is not unlimited. Requests have a 30-second pre-output budget and at most six attempts. Clients should honor Retry-After and use capped exponential backoff with jitter. No 100% uptime guarantee is offered.
GET /health reports configuration readiness, not a live inference guarantee.
JavaScript / Node.js
No SDK required. Node.js 20+ supports this example. Set DEDICATED_API_KEY=abhiaikey in your server environment. Keep production application authentication on your own server.
const response = await fetch(
'https://dedicated-ai.pages.dev/v1/chat/completions',
{
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.DEDICATED_API_KEY}`,
'Content-Type': 'application/json',
},
signal: AbortSignal.timeout(35000),
body: JSON.stringify({
model: 'gpt-6.1-sol',
messages: [
{ role: 'system', content: 'Answer concisely.' },
{ role: 'user', content: 'Write a welcome message.' },
],
max_tokens: 512,
temperature: 0.7,
}),
},
);
const data = await response.json();
if (!response.ok) {
throw new Error(`${response.status}: ${data.error?.message}`);
}
console.log(data.choices[0].message.content);
console.log('Request ID:', response.headers.get('X-Request-Id'));Python
Python 3.10+; no additional packages required. Export DEDICATED_API_KEY=abhiaikey before running.
import json
import os
import urllib.request
import urllib.error
payload = {
"model": "gpt-6.1-sol",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 512,
}
request = urllib.request.Request(
"https://dedicated-ai.pages.dev/v1/chat/completions",
data=json.dumps(payload).encode("utf-8"),
headers={
"Authorization": "Bearer " + os.environ["DEDICATED_API_KEY"],
"Content-Type": "application/json",
},
method="POST",
)
try:
with urllib.request.urlopen(request, timeout=35) as response:
result = json.load(response)
print(result["choices"][0]["message"]["content"])
except urllib.error.HTTPError as error:
print("HTTP", error.code, error.read().decode("utf-8"))
except urllib.error.URLError as error:
print("Connection failed:", error.reason)Streaming client
Use fetch rather than EventSource: this endpoint requires POST and an authorization header. This example handles fragmented UTF-8, multiple events per network chunk, error events and incomplete responses. Pass an AbortController signal to stop generation. Use textContent, not unsanitized HTML, when rendering output.
async function streamChat(messages, key, onText, signal) {
const response = await fetch(
'https://dedicated-ai.pages.dev/v1/chat/completions', {
method: 'POST', signal,
headers: {
Authorization: `Bearer ${key}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'gpt-6.1-sol', messages, stream: true,
max_tokens: 512,
}),
},
);
if (!response.ok) {
const error = await response.json();
throw new Error(error.error?.message || `HTTP ${response.status}`);
}
if (!response.body) throw new Error('Missing stream');
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
try {
while (true) {
const { value, done } = await reader.read();
buffer += done ? decoder.decode()
: decoder.decode(value, { stream: true });
buffer = buffer.replace(/\r\n/g, '\n');
let boundary;
while ((boundary = buffer.indexOf('\n\n')) !== -1) {
const frame = buffer.slice(0, boundary);
buffer = buffer.slice(boundary + 2);
const data = frame.split('\n')
.filter(line => line.startsWith('data:'))
.map(line => line.slice(5).trimStart()).join('\n');
if (!data) continue;
if (data.trim() === '[DONE]') return;
const event = JSON.parse(data);
if (event.error) throw new Error(event.error.message);
const text = event.choices?.[0]?.delta?.content;
if (typeof text === 'string') onText(text);
}
if (done) throw new Error('Stream ended before [DONE]');
}
} finally {
await reader.cancel().catch(() => {});
reader.releaseLock();
}
}
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 125000);
try {
await streamChat(
[{ role: 'user', content: 'Hello!' }],
'abhiaikey', text => console.log(text), controller.signal,
);
} finally { clearTimeout(timer); }
// controller.abort() cancels an active request.Request & response reference
| Field | Default / allowed values |
|---|---|
model | One of the three public IDs; defaults to claud-opus-5 |
messages | Required: 1-100 messages, each with text content and role |
stream | false; boolean only |
max_tokens | 2048; integer 1-8192 |
temperature | Optional number 0-2 |
top_p | Optional number 0-1 |
stop | Optional string or up to four strings |
n | Only 1 supported |
Maximum JSON request size: 256 KB. Combined message text limit: 131,072 characters. This is a text-only chat API; image generation, file uploads, tool calls and structured-output modes are not supported. Conversations are stateless: resend the history needed for each reply.
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1791244800,
"model": "gpt-6.1-sol",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "Hello!"},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 10, "completion_tokens": 3, "total_tokens": 13}
}usage is optional. Read the answer from choices[0].message.content. For streams, concatenate choices[0].delta.content. Keep X-Request-Id for support without logging private prompts.
Production integration
Use a server-side application endpoint to authenticate your users, restrict spending and limit each user's outstanding requests. CORS support is not authentication. Avoid putting private application secrets into browser bundles. The published shared key does not identify individual users.
For short replies, use max_tokens: 256 or 512 and keep history focused. Stream for faster perceived response. Limit concurrent work in your client with a bounded queue and reject or defer excess work rather than growing an unlimited memory queue. Cancel work when your user disconnects.
Retry only 429, 502, 503 or 504 responses, at most twice. Honor Retry-After; add jitter so clients do not retry simultaneously. Do not retry validation/authentication errors. Do not automatically restart a stream after partial output: show the interruption and offer an explicit retry. Requests are not idempotent; retrying a timed-out request may duplicate inference work.
async function chatWithRetry(payload, key) {
const retryable = new Set([429, 502, 503, 504]);
for (let attempt = 0; attempt < 3; attempt++) {
const response = await fetch(
'https://dedicated-ai.pages.dev/v1/chat/completions', {
method: 'POST',
signal: AbortSignal.timeout(35000),
headers: {
Authorization: `Bearer ${key}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({ ...payload, stream: false }),
},
);
if (response.ok) return response.json();
const data = await response.json().catch(() => null);
if (!retryable.has(response.status) || attempt === 2) {
throw new Error(data?.error?.message || `HTTP ${response.status}`);
}
const header = response.headers.get('Retry-After');
const seconds = header ? Number(header) : NaN;
const retryMs = Number.isFinite(seconds) ? seconds * 1000
: header ? Date.parse(header) - Date.now() : 0;
const delay = Math.max(1000 * 2 ** attempt, retryMs || 0);
if (delay > 60000) throw new Error('Retry later; cooldown exceeds 60s');
await new Promise(resolve =>
setTimeout(resolve, delay + Math.random() * 500));
}
}Automatic failover and load-aware routing improve resilience but do not create unlimited inference capacity. A 1,000 requests/second production target requires an appropriate hosting plan, sufficient inference capacity and measured end-to-end load testing. A local mock concurrency test is not proof of live throughput.
Troubleshooting
| Symptom | Check |
|---|---|
| 401 | Use exactly Bearer abhiaikey; remove spaces/newlines in the key |
| 404 | Base URL ends in /v1, not /v1/v1; copy a public model ID |
| 415 | Send Content-Type: application/json |
| 400 | Send a messages array with string content; remove unsupported parameters |
| 429 / 503 | Respect cooldown, reduce concurrency, retry with jitter |
| Incomplete stream | Check SSE error events; do not treat a partial answer as complete |
| Slow first reply | Use streaming and shorter history; failover can take up to 30 seconds before output |
Health check: GET /health needs no key and reports configuration readiness only. Model catalog: GET /v1/models requires authentication. Test an actual chat request separately to verify inference availability.
Privacy
Messages are processed by external services and their data-retention policies apply. Do not send confidential data without checking applicable policies. The frontend stores its access key for the browser session and does not persist chat history. Service credentials stay on the backend, never in the frontend. The shared access key is not per-user authentication: restrict access or add your own authentication before exposing private workloads.