A drop-in OpenAI replacement that runs on your hardware.
LocalAI implements the OpenAI REST surface — chat, embeddings, audio, images — and serves it from local models, so an existing application can be repointed at it by changing a base URL. It brokers several inference backends behind one API and needs no GPU, which makes it a practical way to take an app off a metered provider without rewriting it.
Run open models on your own machine with one command.
LLM inference in plain C/C++, fast enough to run on a laptop.
High-throughput serving built around PagedAttention.
Visual builder for agents and RAG flows, deployable as an API.
Fair-code workflow automation with AI steps built in.
Track the experiments, register the models, ship the good one.
No reviews yet — be the first to review this project.
Sign in to write a review.
Sign In