Call a hosted model in seconds, fine-tune a small model on your own data for lower latency and cost, spin up a dedicated GPU, or run it all in your own AWS account. OpenAI-compatible everywhere.
Get startedCustom fine-tuning
Fine-tune a compact open model (LoRA/QLoRA) on your own data and it can match a much larger model on your specific use case — while responding quicker and costing a fraction as much. Upload your dataset, we train it on a GPU, and you get a private, OpenAI-compatible endpoint.
Start free
Call a ready-to-use open model over an OpenAI-compatible API — no setup, no deployment. Just an API key.
Your model, your GPU
Deploy any Hugging Face model (or a private one with your token) onto its own GPU and get a stable, node-direct HTTPS endpoint.
Bring your own AWS
Run the whole data plane inside your own AWS account — your GPUs, your VPC, your data. We orchestrate; nothing leaves your account.
Shared & Dedicated access is approved per account while we're in early access.