Infrastructure for the agentic era.

Today, every flagship model behind one endpoint, ready for any agent. Next, run and deploy your agents on the same platform. One key. One platform.

Get a key
  • 01Every model · one key
  • 02Zero-downtime serving
  • 03Any agent · drop-in
Runs under every agent you already usedrop-in
Claude CodeCodexCursorGemini CLIGitHub CopilotClineopencodeWindsurfZedWarpReplitJetBrains AI
+ anything that speaks the OpenAI or Anthropic API
The engine[02/07]

One engine handles everything underneath.

You point an agent at an endpoint. We do the rest. You never see the machinery — you see what it guarantees.

  1. 01

    Always on

    Failures never reach your loop. A model that stumbles is replaced mid-request, before your agent notices.

  2. 02

    Fast where agents feel it

    First token, long streams, tool-call turnarounds — tuned for the moments that stall a run.

  3. 03

    Cost engineered down

    The same work costs less here, and keeps getting cheaper. That's the engine's job, not yours.

  4. 04

    Tool calls that hold

    Strict schemas, parallel calls, streams that survive hour-long runs without dropping a frame.

Inside the engine →
FIG. 02 — REQUEST PATHCUTAWAY
The catalog[03/07]

Every model, behind one endpoint.

You hold one key and one endpoint — never another vendor account, quota, or bill. Open-weight flagships sit next to premium frontier models behind the same door, and the catalog keeps growing. One key reaches all of it.

1
endpoint · the whole catalog
4
API protocols · one endpoint
0
changes to your agent
Live in the catalog todaygrowing
Claude Opus 4.8frontierClaude Opus 4.7frontierGPT-5.5frontierGPT-5.4frontierGemini 3 ProfrontierGrok 4.3frontierGPT-OSS 120BworkhorseSonnet 4.6premiumClaude Sonnet 5premiumClaude Haiku 4.5premiumGLM 5.2coding flagshipGLM 5coding flagshipQwen3 Coder 480Bcoding flagshipKimi K2.7 Codecoding flagshipQwen3-Maxcoding flagshipQwen3 235BreasoningDeepSeek V3.2reasoningKimi K2 ThinkingreasoningMiniMax M2.5reasoningERNIE X1reasoningHunyuan T1reasoningStep 3reasoningNemotron Super 3 120BworkhorseNemotron 3 120BworkhorseQwen3 Next 80Bworkhorse
Mistral Large 3 675BworkhorseERNIE 5.1workhorseHunyuan TurboSworkhorseGemini 3 FlashworkhorseQwen3 Coder 30BcodingQwen3 Coder NextcodingKimi K2.5codingDevstral 2 123BcodingGemma 4 26Bcheap · big ctxLongCat 2.0cheap · big ctxGPT-OSS 20Bcheap · fastGLM 4.7 Flashcheap · fastNemotron Nano 3 30Bcheap · fasttext-embedding-3-largeembeddingstext-embedding-3-smallembeddingsvoyage-3.5embeddingsvoyage-code-3embeddingsCohere Embed v4embeddingsGemini EmbeddingembeddingsQwen3-Embedding 8BembeddingsBGE-M3embeddingsjina-embeddings-v3embeddingsMistral EmbedembeddingsNomic Embed v1.5embeddings
Browse every model →
For agents[04/07]

Built for the loop, not the demo.

Chat traffic is easy. Agent traffic is long, tool-heavy, and unforgiving. The serving layer is shaped around that from the start.

[ A ]

Drop-in, both dialects

Speaks the Anthropic and OpenAI APIs. Point your tool at the endpoint — no SDK, no rewrite.

[ B ]

Streams that don't die

Hour-long runs, heavy tool use — the stream holds, or is rescued before your agent ever sees a gap.

[ C ]

Tool calls, strict

Schemas enforced, parallel calls handled, arguments intact on every frame.

[ D ]

Context that holds

A run is one conversation, not a hundred requests. The engine treats it that way.

One command[05/07]

Point any agent here in one command.

Run npx layerx1 and pick your tools — the installer writes each config where the tool expects it, key and model included; GUI-configured tools get their exact paste-in values. No SDK, no code changes. The session below is the setup.

npx layerx1live session · replay
$ npx layerx1 Layer X1 · configure your coding agent Press enter to accept the [default]. Gateway URL: [https://api.layerx1.in] API key (lx1_…): lx1_•••••••••••• Default model: [lx1-gpt-oss-120b] Which tools? (comma-separated numbers, or 'all') 1. Claude Code · Anthropic /v1/messages 2. OpenAI Codex CLI · OpenAI Responses /v1/responses 3. Aider · OpenAI /v1/chat/completions (via LiteLLM) 4. Continue (VS Code / JetBrains) · OpenAI /v1/chat/completions 5. Cline (VS Code) · OpenAI /v1/chat/completions 6. Cursor · OpenAI /v1/chat/completions 7. Windsurf · OpenAI /v1/chat/completions Selection: [all] 1,2 Configuring → https://api.layerx1.in model=lx1-gpt-oss-120b ✓ Claude Code → wrote ~/.claude/settings.json • If you previously logged in to Claude, run `/logout` once so the gateway key is used. • Then just run `claude`. ✓ OpenAI Codex CLI → wrote ~/.codex/config.toml • Persist your key: export LAYERX1_API_KEY in your shell profile (setup offers to do this for you). • Codex requires the Responses API — Layer X1 serves /v1/responses. • Then run `codex`. Persist LAYERX1_API_KEY to your shell profile now? (Y/n) [Y] ✓ LAYERX1_API_KEY added to ~/.zshrc — restart your shell (or source it) to pick it up. Done. Run `npx layerx1 test` to verify the gateway.
Pricing[06/07]

Start free. Scale when your agents do.

Free
$0 /mo
$5 included usage
  • The whole catalog, zero commitment
  • All four API protocols
  • No card required
Starter
$5 /mo
$200 included usage
  • Every model in the catalog
  • All four API protocols
  • Keys in minutes, no call
Pro — recommended
$19 /mo
$600 included usage
  • Everything in Starter
  • Higher rate & concurrency ceilings
  • Built for all-day agent loops
Max
$49 /mo
$3,000 included usage
  • Everything in Pro
  • The heaviest loops, parallel fleets
  • Top rate limits on the platform

Every plan starts free — no card, keys in minutes.

Full pricing & FAQ →
Proof of motion[07/07]

Shipped, not promised.

Changelog — latestall →
2026·07·07Setup guides for 12 more agent harnessesopencode, Crush, Goose, Hermes, OpenClaw, Qwen Code, OpenHands, Zed, Aider, Continue, Cline and Kilo Code — each with verified, copy-paste configuration in the docs.
2026·07·06Serving engine v2Higher burst headroom and steadier serving under peak load, across every model in the catalog.
2026·07·04Streams that surviveIf a stream dies before the first token arrives, the engine fails over automatically — your agent never sees the gap.
2026·07·04Paid plans openStarter $5, Pro $19 and Max $49 subscriptions are live — upgrade from the dashboard, effective immediately.
ManifestoN→∞
“Inference is the first layer, not the last.
Put your agent on real infrastructure

Your agent doesn't change.
Everything underneath does.

$export ANTHROPIC_BASE_URL=https://api.layerx1.in

Start free · no card · Starter from $5/mo