Run open models on your own machine with one command.
Ollama reduced local model hosting to `ollama run`. It bundles weights, quantisation and a chat template into a single pullable artefact, exposes an OpenAI-compatible HTTP API on localhost, and handles GPU offload automatically. It is the shortest path from nothing installed to a working local model, which is why so many other projects list it as their default backend.
LLM inference in plain C/C++, fast enough to run on a laptop.
High-throughput serving built around PagedAttention.
A drop-in OpenAI replacement that runs on your hardware.
Visual builder for agents and RAG flows, deployable as an API.
Fair-code workflow automation with AI steps built in.
Track the experiments, register the models, ship the good one.
No reviews yet — be the first to review this project.
Sign in to write a review.
Sign In