Use Xantly with the OpenAI TypeScript SDK

The openai npm package accepts a baseURL option. Point it at Xantly for smart routing, semantic cache, streaming, and tool calling, works in Node, Bun, Deno, and Edge.

The openai npm package accepts a baseURL constructor option. Point it at Xantly once and chat, streaming, tools, and structured outputs all route through smart routing + cache + memory. Works identically in Node, Bun, Deno, and the Vercel Edge runtime.

Prerequisites

Setup

import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://api.xantly.com/v1',
  apiKey: process.env.XANTLY_API_KEY, // xantly_sk_...
})

That's it. Every existing OpenAI call keeps working, client.chat.completions.create(...), client.embeddings.create(...), everything.

Env-var style (no code change at all):

export OPENAI_BASE_URL=https://api.xantly.com/v1
export OPENAI_API_KEY=xantly_sk_...

Then new OpenAI() picks them up.

Chat completions

const resp = await client.chat.completions.create({
  model: 'xantly/auto-quality',
  messages: [{ role: 'user', content: 'Write a bubble sort in TypeScript.' }],
})
console.log(resp.choices[0].message.content)

Streaming iterator

const stream = await client.chat.completions.create({
  model: 'xantly/auto-quality',
  messages: [{ role: 'user', content: 'Stream a haiku.' }],
  stream: true,
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}

Tool calling

const tools = [{
  type: 'function' as const,
  function: {
    name: 'get_weather',
    description: 'Get the weather for a city.',
    parameters: {
      type: 'object',
      properties: { city: { type: 'string' } },
      required: ['city'],
    },
  },
}]

const resp = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-4.6',
  messages: [{ role: 'user', content: 'Weather in Paris?' }],
  tools,
})
const call = resp.choices[0].message.tool_calls?.[0]
console.log(call?.function.name, call?.function.arguments)

Tool-call shape is translated across providers, swap model between OpenAI/Claude/Groq without changing the tool schema.

Structured outputs

import { z } from 'zod'
import { zodResponseFormat } from 'openai/helpers/zod'

const User = z.object({
  name: z.string(),
  email: z.string(),
  role: z.string(),
})

const resp = await client.beta.chat.completions.parse({
  model: 'openai/gpt-5.4',
  messages: [{ role: 'user', content: 'Extract: Jane Doe, [email protected], CTO' }],
  response_format: zodResponseFormat(User, 'user'),
})

console.log(resp.choices[0].message.parsed)

Vercel AI SDK interop

If you use the Vercel AI SDK, Xantly plugs in via @ai-sdk/openai-compatible instead of this client. See the dedicated Vercel AI SDK integration.

Model choice

Model IDWhen
xantly/auto-qualityBaRP on T1 pool, production default.
xantly/auto-valueBalanced T2.
xantly/auto-speedLowest latency, shortest completions.
openai/gpt-5.4Pin OpenAI.
anthropic/claude-sonnet-4.6Pin Claude, response shape translated.
groq/llama-3.3-70bUltra-fast Groq.

Verify

const resp = await client.chat.completions.create({
  model: 'xantly/auto-speed',
  messages: [{ role: 'user', content: 'say pong' }],
})
console.log(resp.choices[0].message.content)

Open your Xantly dashboard, you'll see the call with routed-to model, cache status, USD cost.

What you get

Gotchas

baseURL (camelCase) not base_url. The TS SDK uses baseURL. Must include /v1.

Edge runtimes and streaming. Cloudflare Workers / Vercel Edge need stream over fetch, the openai SDK handles this automatically. No manual polyfills needed.

Type imports. If you see Parameters<typeof client.chat.completions.create> errors, make sure you're on openai v4.0.0+.

Browser usage. Don't ship xantly_sk_... to the browser. Proxy calls through your own server or use server-side API routes.

Next steps