Dedicated LLMs

DEVELOPER REFERENCE

Dedicated LLMs API

OpenAI-compatible text chat with automatic failover.

Quickstart

Base URL
API key
abhiaikey
Chat completion

POST /v1/chat/completions accepts text system, user and assistant messages. Options: stream, max_tokens (1-8192), temperature (0-2), top_p (0-1), and stop. Only n=1 is supported. Tools, images and structured output are not supported.

Models

These are custom assistant aliases, not verified releases of the similarly named vendor models. The response model field stays on your selected alias while inference may use a different configured provider.

API model IDAssistant name
claud-opus-5claud opus 5
gpt-6.1-solgpt 6.1 sol
glm-5.2glm 5.2

GET /v1/models returns exactly these three aliases. All /v1 requests require Authorization: Bearer abhiaikey or X-API-Key: abhiaikey.

Streaming

Set "stream": true to receive real server-sent events ending in data: [DONE]. Provider failover is allowed before the first output token. After output begins, interruption is reported as an SSE error instead of mixing answers from different models.

data: {"object":"chat.completion.chunk","model":"gpt-6.1-sol","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: [DONE]

Errors & limits

StatusMeaning
400 / 413 / 415Invalid, oversized or unsupported request
401Invalid Dedicated LLMs key
404Unknown model alias or endpoint
429Temporary rate limit
503Service temporarily unavailable

Requests scale across the service network, but processing capacity and upstream quotas still apply. Free inference is not unlimited. Requests have a 30-second pre-output budget and at most six attempts. Clients should honor Retry-After and use capped exponential backoff with jitter. No 100% uptime guarantee is offered.

GET /health reports configuration readiness, not a live inference guarantee.

JavaScript / Node.js

No SDK required. Node.js 20+ supports this example. Set DEDICATED_API_KEY=abhiaikey in your server environment. Keep production application authentication on your own server.

Server-side chat
const response = await fetch(
  'https://dedicated-ai.pages.dev/v1/chat/completions',
  {
    method: 'POST',
    headers: {
      Authorization: `Bearer ${process.env.DEDICATED_API_KEY}`,
      'Content-Type': 'application/json',
    },
    signal: AbortSignal.timeout(35000),
    body: JSON.stringify({
      model: 'gpt-6.1-sol',
      messages: [
        { role: 'system', content: 'Answer concisely.' },
        { role: 'user', content: 'Write a welcome message.' },
      ],
      max_tokens: 512,
      temperature: 0.7,
    }),
  },
);
const data = await response.json();
if (!response.ok) {
  throw new Error(`${response.status}: ${data.error?.message}`);
}
console.log(data.choices[0].message.content);
console.log('Request ID:', response.headers.get('X-Request-Id'));

Python

Python 3.10+; no additional packages required. Export DEDICATED_API_KEY=abhiaikey before running.

Standard-library client
import json
import os
import urllib.request
import urllib.error

payload = {
    "model": "gpt-6.1-sol",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 512,
}
request = urllib.request.Request(
    "https://dedicated-ai.pages.dev/v1/chat/completions",
    data=json.dumps(payload).encode("utf-8"),
    headers={
        "Authorization": "Bearer " + os.environ["DEDICATED_API_KEY"],
        "Content-Type": "application/json",
    },
    method="POST",
)
try:
    with urllib.request.urlopen(request, timeout=35) as response:
        result = json.load(response)
        print(result["choices"][0]["message"]["content"])
except urllib.error.HTTPError as error:
    print("HTTP", error.code, error.read().decode("utf-8"))
except urllib.error.URLError as error:
    print("Connection failed:", error.reason)

Streaming client

Use fetch rather than EventSource: this endpoint requires POST and an authorization header. This example handles fragmented UTF-8, multiple events per network chunk, error events and incomplete responses. Pass an AbortController signal to stop generation. Use textContent, not unsanitized HTML, when rendering output.

Browser or Node.js streaming
async function streamChat(messages, key, onText, signal) {
  const response = await fetch(
    'https://dedicated-ai.pages.dev/v1/chat/completions', {
      method: 'POST', signal,
      headers: {
        Authorization: `Bearer ${key}`,
        'Content-Type': 'application/json',
      },
      body: JSON.stringify({
        model: 'gpt-6.1-sol', messages, stream: true,
        max_tokens: 512,
      }),
    },
  );
  if (!response.ok) {
    const error = await response.json();
    throw new Error(error.error?.message || `HTTP ${response.status}`);
  }
  if (!response.body) throw new Error('Missing stream');
  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  try {
    while (true) {
      const { value, done } = await reader.read();
      buffer += done ? decoder.decode()
        : decoder.decode(value, { stream: true });
      buffer = buffer.replace(/\r\n/g, '\n');
      let boundary;
      while ((boundary = buffer.indexOf('\n\n')) !== -1) {
        const frame = buffer.slice(0, boundary);
        buffer = buffer.slice(boundary + 2);
        const data = frame.split('\n')
          .filter(line => line.startsWith('data:'))
          .map(line => line.slice(5).trimStart()).join('\n');
        if (!data) continue;
        if (data.trim() === '[DONE]') return;
        const event = JSON.parse(data);
        if (event.error) throw new Error(event.error.message);
        const text = event.choices?.[0]?.delta?.content;
        if (typeof text === 'string') onText(text);
      }
      if (done) throw new Error('Stream ended before [DONE]');
    }
  } finally {
    await reader.cancel().catch(() => {});
    reader.releaseLock();
  }
}

const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 125000);
try {
  await streamChat(
    [{ role: 'user', content: 'Hello!' }],
    'abhiaikey', text => console.log(text), controller.signal,
  );
} finally { clearTimeout(timer); }
// controller.abort() cancels an active request.

Request & response reference

FieldDefault / allowed values
modelOne of the three public IDs; defaults to claud-opus-5
messagesRequired: 1-100 messages, each with text content and role
streamfalse; boolean only
max_tokens2048; integer 1-8192
temperatureOptional number 0-2
top_pOptional number 0-1
stopOptional string or up to four strings
nOnly 1 supported

Maximum JSON request size: 256 KB. Combined message text limit: 131,072 characters. This is a text-only chat API; image generation, file uploads, tool calls and structured-output modes are not supported. Conversations are stateless: resend the history needed for each reply.

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1791244800,
  "model": "gpt-6.1-sol",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "Hello!"},
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 10, "completion_tokens": 3, "total_tokens": 13}
}

usage is optional. Read the answer from choices[0].message.content. For streams, concatenate choices[0].delta.content. Keep X-Request-Id for support without logging private prompts.

Production integration

Use a server-side application endpoint to authenticate your users, restrict spending and limit each user's outstanding requests. CORS support is not authentication. Avoid putting private application secrets into browser bundles. The published shared key does not identify individual users.

For short replies, use max_tokens: 256 or 512 and keep history focused. Stream for faster perceived response. Limit concurrent work in your client with a bounded queue and reject or defer excess work rather than growing an unlimited memory queue. Cancel work when your user disconnects.

Retry only 429, 502, 503 or 504 responses, at most twice. Honor Retry-After; add jitter so clients do not retry simultaneously. Do not retry validation/authentication errors. Do not automatically restart a stream after partial output: show the interruption and offer an explicit retry. Requests are not idempotent; retrying a timed-out request may duplicate inference work.

Bounded retries for non-streaming requests
async function chatWithRetry(payload, key) {
  const retryable = new Set([429, 502, 503, 504]);
  for (let attempt = 0; attempt < 3; attempt++) {
    const response = await fetch(
      'https://dedicated-ai.pages.dev/v1/chat/completions', {
        method: 'POST',
        signal: AbortSignal.timeout(35000),
        headers: {
          Authorization: `Bearer ${key}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ ...payload, stream: false }),
      },
    );
    if (response.ok) return response.json();
    const data = await response.json().catch(() => null);
    if (!retryable.has(response.status) || attempt === 2) {
      throw new Error(data?.error?.message || `HTTP ${response.status}`);
    }
    const header = response.headers.get('Retry-After');
    const seconds = header ? Number(header) : NaN;
    const retryMs = Number.isFinite(seconds) ? seconds * 1000
      : header ? Date.parse(header) - Date.now() : 0;
    const delay = Math.max(1000 * 2 ** attempt, retryMs || 0);
    if (delay > 60000) throw new Error('Retry later; cooldown exceeds 60s');
    await new Promise(resolve =>
      setTimeout(resolve, delay + Math.random() * 500));
  }
}

Automatic failover and load-aware routing improve resilience but do not create unlimited inference capacity. A 1,000 requests/second production target requires an appropriate hosting plan, sufficient inference capacity and measured end-to-end load testing. A local mock concurrency test is not proof of live throughput.

Troubleshooting

SymptomCheck
401Use exactly Bearer abhiaikey; remove spaces/newlines in the key
404Base URL ends in /v1, not /v1/v1; copy a public model ID
415Send Content-Type: application/json
400Send a messages array with string content; remove unsupported parameters
429 / 503Respect cooldown, reduce concurrency, retry with jitter
Incomplete streamCheck SSE error events; do not treat a partial answer as complete
Slow first replyUse streaming and shorter history; failover can take up to 30 seconds before output

Health check: GET /health needs no key and reports configuration readiness only. Model catalog: GET /v1/models requires authentication. Test an actual chat request separately to verify inference availability.

Privacy

Messages are processed by external services and their data-retention policies apply. Do not send confidential data without checking applicable policies. The frontend stores its access key for the browser session and does not persist chat history. Service credentials stay on the backend, never in the frontend. The shared access key is not per-user authentication: restrict access or add your own authentication before exposing private workloads.