OpenAI-compatible ยท Local inference
One API for your
local Llama models
Route requests to Llama running on your own hardware. Drop-in replacement for Groq and OpenAI โ same SDK, your server.
Ollama online
ยท
2 models available
OpenAI-compatible
Use the same SDKs and libraries. Change only the base URL and API key.
Your infrastructure
Data stays on your server. No third-party API calls or per-token billing.
Multiple models
Switch between Llama models for quality vs speed โ all through one endpoint.
Available models
llama3.2:3b
Local
llama3.1:8b
Local
Ready to build?
Create an account, generate an API key, and start calling your local Llama in minutes.