Inference for AI agents.

An LLM API for coding agents. Run Claude Code cheaper on GLM, MiniMax, Kimi — or Codex, Cline, OpenCode — through one OpenAI-compatible endpoint, built for long-running agent loops. One key, one bill.

  • 01Sixteen harnesses · one command
  • 02Four protocols · one key
  • 03Fails over mid-stream
from openai import OpenAI

client = OpenAI(
    base_url="https://api.layerx1.com/v1",
    api_key="lx1_your_key",
)

response = client.chat.completions.create(
    model="lx1-gpt-oss-120b",
    messages=[{"role": "user", "content": "Say hello."}],
)
Both dialects · one lx1_ key · every modelFull API reference →
One command[01/05]

Point any agent here in one command.

Run npx layerx1 and pick your tools — the installer writes each config where the tool expects it, key and model included; GUI-configured tools get their exact paste-in values. No SDK, no code changes. The session below is the setup.

Every harness we write configs for[16]
Claude CodeCodex CLICursoropencodeOpenClawHermesGooseClineAiderContinueWindsurfCrushQwen CodeOpenHandsZedKilo Code
npx layerx1live session · replay
$ npx layerx1 ██╗ █████╗ ██╗ ██╗███████╗██████╗ ██╗ ██╗ ██╗ ██║ ██╔══██╗╚██╗ ██╔╝██╔════╝██╔══██╗ ╚██╗██╔╝███║ ██║ ███████║ ╚████╔╝ █████╗ ██████╔╝ ╚███╔╝ ╚██║ ██║ ██╔══██║ ╚██╔╝ ██╔══╝ ██╔══██╗ ██╔██╗ ██║ ███████╗██║ ██║ ██║ ███████╗██║ ██║ ██╔╝ ██╗ ██║ ╚══════╝╚═╝ ╚═╝ ╚═╝ ╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═╝ ╚═╝ ▚▚▚ Layer X1 · one endpoint · one key · every model ──────────────────────────────────────────────────────── step 1/5 Pick your tools space toggles · already-installed tools are pre-selected Which tools should point at Layer X1? 2 already installed · space toggle · a all · enter confirm · esc cancel Terminal agents ❯ ◉ Claude Code installed · writes config ◉ OpenAI Codex CLI installed · writes config ◯ Aider not installed · writes config ◯ Hermes Agent not installed · writes config Editors & extensions ◯ Continue (VS Code / JetBrains) not installed · writes config 2 selected step 2/5 Install what's missing installed tools are skipped ✓ Claude Code already installed — skipped ✓ OpenAI Codex CLI already installed — skipped ✓ everything you picked is already installed. step 3/5 Choose your models loaded live from the gateway — new models appear here automatically ✓ 34 models live from https://api.layerx1.com Primary model the default your agent uses · recommended: auto Coding flagship ○ lx1-qwen3-coder-480b 131k ctx · $0.45/$1.8 per Mtok ○ lx1-glm-5.2 131k ctx · $1.4/$4.4 per Mtok · reasoning Workhorse ❯ ● auto plan-aware route · GLM Flash default · reasoning step 4/5 Your API key pasted input is masked and only written to the tool configs you picked Paste your Layer X1 API key: lx1_•••••••••••• step 5/5 Review and apply gateway https://api.layerx1.com primary model auto small / fast lx1-glm-4.7-flash key lx1_YO••••••••••KEY tools Claude Code, OpenAI Codex CLI Write these configs? [Y/n] yes ✓ Claude Code → wrote ~/.claude/settings.json • If you previously logged in to Claude, run `/logout` once so the gateway key is used. • Then just run `claude`. ✓ OpenAI Codex CLI → wrote ~/.codex/config.toml • The API key is in the provider block — Codex Desktop does not need LAYERX1_API_KEY. • Codex requires the Responses API — Layer X1 serves /v1/responses. • Then run `codex`. ✓ gateway replied in 412ms · served by auto ──────────────────────────────────────────────────────── You're set. → Claude Code claude → OpenAI Codex CLI codex → check anytime npx layerx1 status
GPT-6 AstrafrontierGPT-5.6 SolfrontierGPT-5.6 TerrafrontierGPT-5.6 LunafrontierMAI Thinking 1reasoningGLM 5.3codingClaude Opus 5frontierClaude Fable 5.1frontierClaude Fable 5frontierClaude Opus 4.7frontierGPT-5.5frontierGPT-5.4frontierGemini 3 ProfrontierKimi K3frontierDeepSeek V4 ProcodingQwen3.7 MaxfrontierInklingfrontierGPT-OSS 120Bgeneral purposeQwen 3.8 27Bgeneral purposeSonnet 4.6premiumClaude Sonnet 5premiumClaude Haiku 4.5premiumGrok 4.3premiumGLM 5.3 FlashcodingGLM 5.2codingGLM 5codingQwen3 Coder 480BcodingKimi K2.7 CodecodingQwen3-MaxcodingQwen3 235BreasoningDeepSeek V3.2reasoningKimi K2 ThinkingreasoningMiniMax M2.5reasoningERNIE X1reasoningHunyuan T1reasoningStep 3reasoningQwen3.5 397BreasoningNemotron 3 UltrareasoningNemotron Super 3 120Bgeneral purposeNemotron 3 120Bgeneral purposeQwen3 Next 80Bgeneral purposeGemma 4 31Bgeneral purposeMistral Large 3 675Bgeneral purposeERNIE 5.1general purposeHunyuan TurboSgeneral purposeGemini 3 Flashgeneral purposeMiniMax M3general purposeKimi K2.6general purposeGLM 4.7general purposeQwen3.7 Plusgeneral purposeQwen3 VL 235Bgeneral purposeHunyuan Hy3general purposeLlama 3.3 70Bgeneral purposeQwen3 Coder 30Beveryday codingQwen3 Coder Nexteveryday coding
What the endpoint promises[02/05]

Three guarantees, drawn to scale.

You point an agent at an endpoint. We do the rest. You never see the machinery — you see what it guarantees. Open one for the full story.

The engine[03/05]

One engine underneath. Every dial in your hands.

You never see the machinery — you see what it guarantees. And you see everything it does for you: usage, spend, every request, caps and keys, live in the dashboard from the first call.

request pathsealed
Spend caps
Request logs
Keys
Usage & spend
Every request accounted — usage · logs · caps · keysOpen your dashboard →
Integrations[04/05]

Wired into the agents you already run.

Six of the harnesses people point here every day — open one for the exact config. Each is a base-URL swap and a key: no SDK, no code changes, and the model picker fills from the catalog.

Ten more with exact copy-paste configs — every supported agent →[16]
Or the agent you're building

Anything that speaks the OpenAI or Anthropic API speaks to us.

Point your SDK at the endpoint and every model in the catalog is a string away — streaming, tool calls, and structured output included. Free tier on the same key.

Plans[05/05]
Free$2 credit to start
Launch
$2 /mo
$30 of usage included

Real monthly headroom at the smallest possible commitment.

  • $30 of model usage included every month
  • Open models + GPT-5.6 Luna
  • 100 requests / 5 min · 5 concurrent
  • In-flight runs never cut off
Starter
$5 /mo
$60 of usage included

A serious month of agent work for the price of a coffee.

  • $60 of model usage included every month
  • Adds GPT-5.6 Terra & the Sonnet class
  • 300 requests / min · 10 concurrent
  • In-flight runs never cut off
ProRecommended
$19 /mo
$300 of usage included

For a daily driver — coding agents that work all day.

  • $300 of model usage included every month
  • Adds GPT-5.6 Sol & the Opus class
  • 1,000 requests / min · 40 concurrent
  • In-flight runs never cut off
Max
$49 /mo
$800 of usage included

5x Pro. For agents that never sleep and teams of one that ship like ten.

  • $800 of model usage included every month
  • Every model — including Astra & Fable
  • 3,000 requests / min · 100 concurrent
  • In-flight runs never cut off
Pay as you go

No subscription. Load credit and spend it whenever you like — at each model's published rate, zero markup. Credit doesn't expire and doesn't reset, and it keeps working after a monthly pool runs out.

Enterprise

Need more than Max — higher limits, custom usage pools, procurement? Talk to us and we'll size a plan to your fleet.

The dashboard and allowance headers show what each request deducts. Model-page prices are the PAYG/list-price reference — see Plans & limits for how the meter works.

Ready when you are

Every agent you run.
One key, one endpoint.

One command wires Claude Code, Codex, Cursor and a dozen more to the same gateway — no per-tool accounts, no per-tool bills.

Start free · no card · Launch from $2/mo